Speech, voice and audio

    Text-to-speech

    Also called TTS, Speech synthesis.

    Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.

    Modern systems do not assemble recorded fragments. They generate the waveform, which is why they can produce a sentence nobody ever said, in a voice that never said it, with breath and hesitation in roughly the right places. It also means output is stochastic: run the same script twice and you get two different readings, differing in pace and emphasis rather than in words.

    That variability is a feature if you treat it as casting. Generating a couple of takes and picking is faster than fighting one take, and it is how most people get a delivery they are happy with without touching a single control.

    The controls that matter are usually the voice itself, the language, and the delivery direction — either as separate settings or as markup inside the text, depending on the provider. Speed is the exception worth knowing about: adjusting it at generation time sounds better than time-stretching the file afterwards.

    In practice

    • Punctuation is prosody. Commas and full stops do more for pacing than any slider.
    • Write numbers, dates and acronyms the way you want them read aloud.
    • Two takes of the same script are two performances — audition rather than re-roll.

    Text-to-audio models

    Catalog entries that turn written text into speech or sound. 19 of the 296 models in the Versely catalog qualify.

    ModelProviderType
    Gemini 3.1 Flash TTSGoogleAudio
    Cartesia Sonic 3.5CartesiaAudio
    Inworld TTS 1.5 MaxInworldAudio
    Inworld TTS 2InworldAudio
    ElevenLabs MultilingualKIEAudio
    Qwen 3 TTS 0.6BQwenAudio
    Seed Audio 1.0ByteDanceAudio
    Suno Sounds V5.5SunoAudio

    Browse all 11 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Feeding in a paragraph written to be read silently. Sentences that scan fine on a page are frequently too long to say in one breath, and synthesis follows your punctuation faithfully.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.