Cartesia Sonic 3.5: Sub-100ms Voices and What Latency Buys You
Cartesia's Sonic 3.5 hits sub-90ms latency across 42 languages and fixes the reads that break TTS: numbers, codes, emails. What that actually buys creators.
Nobody watching a finished voiceover cares that it rendered in 90 milliseconds instead of 900. They care whether the read is good. So when a TTS provider leads its launch with a latency number, the first honest question is: latency for whom?
Cartesia's answer with Sonic 3.5 is mostly "for live voice agents" — but the model shipped alongside real gains that matter for anyone generating rendered voiceover too. Here's what actually changed, where the speed genuinely pays off, and where it's mostly beside the point.
What Cartesia shipped
Cartesia launched Sonic 3.5 alongside Ink-2, its speech-to-text model, positioning the pair as a unified real-time voice stack rather than two separate products. That framing matters: this is a launch built for the round trip of a live conversation — hear the user, understand them, respond — not just for reading a script out loud.
The number: under 90 milliseconds
Per Cartesia's own model documentation, Sonic 3.5 delivers sub-90ms first-audio latency — the time between sending text and hearing the first audible sample come back — with native support across 42 languages. First-audio latency is the specific metric that matters for turn-taking: it's the gap a listener actually perceives as "did this thing hear me and start responding," not the total time to generate an entire clip.
The same documentation calls out something less flashy but arguably more useful day to day: stronger handling of the text that reliably breaks naive TTS — numbers, phone numbers, email addresses, codes and heteronyms (words spelled the same but pronounced differently depending on context, like the noun and verb forms of "record"). Getting those wrong is the fastest way to make a synthetic voice sound synthetic, and fixing them normally means manually rewriting your script into phonetic spellings before you generate anything. Sonic 3.5's pitch is that you shouldn't have to.
42 languages, natively — not translated after the fact
"Native support" is doing real work in that spec. A model that's fluent in one language and merely functional in forty-one others behaves differently from one trained across all of them from the start — accent, prosody and the handling of language-specific quirks (numbers read the German way versus the English way, for instance) tend to be where the difference shows up first. For a creator producing the same ad or explainer in several markets, that's the practical payoff of "native": fewer awkward line readings to catch and re-generate per language, not just a longer dropdown of voice options.
Who actually needs sub-90ms
The tell for what this spec is really built for: LiveKit — infrastructure specifically for real-time voice agents — highlighted day-one access to Sonic 3.5 as a launch-day partner. That's not a coincidence. Below a certain round-trip latency threshold, a voice agent stops feeling like it's lagging or interrupting and starts feeling like a conversation. Above it, every exchange has a small but noticeable hitch that users register as "the AI thinking" even when the underlying reasoning was instant — the bottleneck was purely how fast the voice came back.
For that use case, shaving latency from the 200–300ms range down to sub-90ms is a genuine product difference, not a marketing rounding error.
Where latency buys you less: rendered voiceover
Here's the honest part. If you're generating a voiceover for a video — the far more common job for most creators — you're not in a live turn-taking loop. You send a script, you get a file back, and that file gets baked into an export that a viewer watches later. The viewer never experiences the 90 milliseconds versus 900 milliseconds gap at all; that variable lives entirely on the production side, before anything ships.
Where speed still helps in that world is iteration, not playback. Generating a couple of takes and auditioning them is a faster way to land a good read than fighting one slow take — and a genuinely fast underlying model means you can burn through more script variants, more voice options and more delivery styles in the same working session. That's a real workflow benefit. It's just a different benefit than "the app feels instant," and worth naming as such rather than importing a live-agent spec sheet wholesale into a content pipeline where it doesn't fully apply.
Sonic 3.5 versus a more expressive model
Speed and technical-text robustness aren't the only axis that matters when picking a TTS model. Inworld TTS 2 sits at the other end of the trade-off: 100+ languages and natural-language style steering, built for fine-grained control over delivery rather than for round-trip speed. The rough decision rule:
- Pick Sonic 3.5 for scripted content with a lot of specifics — ad reads with prices and phone numbers, product explainers with model codes, multilingual variants of the same script — where broad language coverage and clean reads on technical text matter more than nuanced emotional direction.
- Pick Inworld TTS 2 when the performance itself is the product — narrative voiceover, character work, anything where you're steering delivery with natural-language instructions rather than just picking a voice and hitting generate.
Neither is strictly better; they're built for different jobs on either side of the same TTS category.
A concrete walkthrough: fast VO iteration in Versely
This is the workflow Sonic 3.5's speed actually earns its keep in — auditioning takes quickly rather than committing to the first one:
1. Write the script with punctuation doing the pacing work. Commas and full stops control the read more than any setting does, and numbers, dates and acronyms should be written the way you want them spoken aloud.
2. Ask the agent for a few takes at once. Something like: "Generate this script in Cartesia Sonic 3.5, warm and conversational, and give me two takes so I can pick." Under the hood this is the generate_speech tool, which takes a model and a voice_id — Sonic 3.5 plus a stock or cloned voice.
3. Attach the winner to your video. Once you've picked a take, attach_audio_to_video lays it onto the footage — mode replace if the narration should be the only audio, or mode mix to layer it over existing ambient sound. This is the same tool used in Versely's add-voiceover-to-video workflow.
4. Re-render only what changed. If one line reads wrong, regenerate that line rather than the whole script — the fast model makes single-line reshoots cheap enough to actually do instead of living with a line that's almost right.
FAQ
What is Cartesia Sonic 3.5?
Sonic 3.5 is Cartesia's newest text-to-speech model, launched alongside the Ink-2 speech-to-text model as a combined real-time voice stack. It delivers sub-90ms first-audio latency across 42 languages, with improved handling of numbers, codes, emails and heteronyms.
How fast is Sonic 3.5, really?
Cartesia's own documentation states sub-90ms first-audio latency — the time from sending text to hearing the first audible sample. That figure is specifically meaningful for live, turn-taking voice agents rather than for a voiceover you generate once and export into a finished video.
Does low latency matter for pre-recorded voiceover?
Less directly than you'd expect. A viewer never perceives the generation-time latency of a finished export. Where speed still helps is iteration — auditioning multiple takes quickly during production, rather than the final playback experience.
Sonic 3.5 or Inworld TTS 2 — which should I use?
Sonic 3.5 for technical, specific scripts across many languages, where clean reads of numbers and codes matter. Inworld TTS 2 when the delivery itself needs fine-grained natural-language style direction, such as narrative or character voiceover.
Sonic 3.5 is live in Versely's AI text-to-speech tool — or browse the full voice-over stack if you're deciding between models rather than committed to one.