Guides

    Cartesia Sonic 3.5: a voice file, not a talking generate (5cr)

    Cartesia Sonic 3.5 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team4 min read

    Cartesia Sonic 3.5 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Cartesia Sonic 3.5 is Cartesia's newest TTS model. The catalog description: multilingual, highly expressive, with emotion control, speed tuning, and voice cloning. Categories: text-to-audio and voice-clone. Content type is audio. It does not require an image. No duration chips. No resolution. Display price: 5 credits. This is a wav (and a clone, if you brought a sample). It is not a face.

    Clone and emotion are still a file

    Voice cloning is the feature that fools people into thinking they have a presenter. You have a speaker. You do not have a jaw. Emotion control and speed tuning change the performance of that speaker: warmer, faster, a held vowel. The output is still text-to-audio.

    The clone category is why this row is not a name-swap with a TTS-only engine. Best voice cloning model is the ranked clone list. Best text-to-speech model is the speech list. Sonic 3.5 can appear in both conversations because it carries both categories. Neither list is a talking-head list.

    The door is AI text-to-speech. The clone path is AI voice cloning. Use them for the track. Do not use them as a shortcut around a plate.

    Five credits, two categories, one job

    5 credits is the display price. Billing on the row is per 1,000 characters; the number you plan around is 5 credits for the speech job. There is no 720p because there is no frame. There is no 5s chip because length is the script, not a video duration.

    Multilingual is not lipsync. A second language is a second wav. If each language needs a mouth, you still need a plate and a lipsync pass per language. Sonic 3.5 will not grow a face because the script switched codes.

    Audio tags and emotion controls punch delivery. They do not open a mouth. If the take goes flat, chunk the script — same problem as narration that flattens.

    When we see the mouth, leave Sonic

    A closed mouth with a cloned voice under it is the production tell. You paid 5 credits for a file and hoped the still would act. The acting row is lipsync: plate plus track in, talking picture out.

    Kling Avatar Pro is one such row (image required, 4K, 58 credits). Run it from AI lipsync. Generate or clone the voice on Sonic 3.5 first. Lock the still on an image row. Then do the mouth. Style-lock the voice across takes so the lipsync pass is not chasing a different speaker every line.

    A cloned speaker is still not a presenter

    The test is visual. If the cut shows teeth, this generate is the wrong last step. If the cut is VO over b-roll, 5 credits and you are done. Sonic 3.5 is "newest and top-ranked" in its own description as a TTS model — that ranking is a speech ranking. It is not a reason to skip the picture.

    FAQ

    Does Cartesia Sonic 3.5 generate a talking avatar?

    No. It writes audio. Categories are text-to-audio and voice-clone. 5 credits, no image, no resolution. The mouth is a lipsync row.

    What does voice cloning change?

    It changes whose voice the file is. Emotion and speed change how that voice performs. None of that tracks lips on a still. Clone here, then lipsync elsewhere.

    Why would I pick Sonic 3.5 over a talking-head model?

    When you need the track: multilingual VO, a cloned speaker, emotion and speed control, 5 credits. When you need the face, you pick a row that takes a plate.

    Can I drop this wav on a product still and ship it?

    As VO, yes. As a presenter, no — the mouth will not move. If we see the mouth, pick lipsync instead of laying Sonic on a closed plate.