Modern systems do not assemble recorded fragments. They generate the waveform, which is why they can produce a sentence nobody ever said, in a voice that never said it, with breath and hesitation in roughly the right places. It also means output is stochastic: run the same script twice and you get two different readings, differing in pace and emphasis rather than in words.
That variability is a feature if you treat it as casting. Generating a couple of takes and picking is faster than fighting one take, and it is how most people get a delivery they are happy with without touching a single control.
The controls that matter are usually the voice itself, the language, and the delivery direction — either as separate settings or as markup inside the text, depending on the provider. Speed is the exception worth knowing about: adjusting it at generation time sounds better than time-stretching the file afterwards.
In practice
- Punctuation is prosody. Commas and full stops do more for pacing than any slider.
- Write numbers, dates and acronyms the way you want them read aloud.
- Two takes of the same script are two performances — audition rather than re-roll.
Text-to-audio models
Catalog entries that turn written text into speech or sound. 22 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Cartesia Sonic 3.6 | Cartesia | Audio |
| Gemini 3.8 Flash TTS | Audio | |
| Inworld TTS 2 | Inworld | Audio |
| Inworld TTS 2 Flash | Inworld | Audio |
| Gemini 3.1 Flash TTS | Audio | |
| Cartesia Sonic 3.5 | Cartesia | Audio |
| MiniMax Speech | MiniMax | Audio |
| ElevenLabs Multilingual | KIE | Audio |
Browse all 13 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Feeding in a paragraph written to be read silently. Sentences that scan fine on a page are frequently too long to say in one breath, and synthesis follows your punctuation faithfully.
Where you will run into it
- Add a Voiceover to a Video — Type the script. Get a narrated video back.
- AI Voice Cloning & Text to Speech — Your voice. Any language. Any script.
Related terms
Voice cloning
Voice cloning meaning: building a reusable synthetic voice from a real sample so new scripts can be spoken in that voice later.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
Voice stability
Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.
AI dubbing
AI dubbing meaning: replacing spoken audio with another language, usually keeping the speaker voice and optionally re-syncing lips.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.