Modern systems do not assemble recorded fragments. They generate the waveform, which is why they can produce a sentence nobody ever said, in a voice that never said it, with breath and hesitation in roughly the right places. It also means output is stochastic: run the same script twice and you get two different readings, differing in pace and emphasis rather than in words.
That variability is a feature if you treat it as casting. Generating a couple of takes and picking is faster than fighting one take, and it is how most people get a delivery they are happy with without touching a single control.
The controls that matter are usually the voice itself, the language, and the delivery direction — either as separate settings or as markup inside the text, depending on the provider. Speed is the exception worth knowing about: adjusting it at generation time sounds better than time-stretching the file afterwards.
In practice
- Punctuation is prosody. Commas and full stops do more for pacing than any slider.
- Write numbers, dates and acronyms the way you want them read aloud.
- Two takes of the same script are two performances — audition rather than re-roll.
Text-to-audio models
Catalog entries that turn written text into speech or sound. 19 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Gemini 3.1 Flash TTS | Audio | |
| Cartesia Sonic 3.5 | Cartesia | Audio |
| Inworld TTS 1.5 Max | Inworld | Audio |
| Inworld TTS 2 | Inworld | Audio |
| ElevenLabs Multilingual | KIE | Audio |
| Qwen 3 TTS 0.6B | Qwen | Audio |
| Seed Audio 1.0 | ByteDance | Audio |
| Suno Sounds V5.5 | Suno | Audio |
Browse all 11 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Feeding in a paragraph written to be read silently. Sentences that scan fine on a page are frequently too long to say in one breath, and synthesis follows your punctuation faithfully.
Where you will run into it
- Add a Voiceover to a Video — Type the script. Get a narrated video back.
- AI Voice Cloning & Text to Speech — Your voice. Any language. Any script.
Related terms
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
Voice stability
Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.