Workflows

    Native audio versus TTS on the same brief

    If the model can speak in-shot, do not generate silent and slap TTS on unless you need a locked brand voice.

    Versely Team6 min read

    If the model can speak in-shot, do not generate silent and slap TTS on unless you need a locked brand voice.

    That is the whole rule. The rest is when to break it, and the one thing you must never do: mix both on one mouth.

    Native audio is picture and soundtrack from the same pass — dialogue, room, footsteps, agreeing by construction. TTS is a voice you control: a cloned founder, a style-locked narrator, the same timbre across 30 episodes. Native audio wins lips in frame. TTS wins a cloned founder voice across a series. They are different contracts. Pick one per mouth.

    Two sounds, two jobs

    Native audio is a video-model feature. Veo 3.1, Seedance 2.5, Happy Horse, and a short list of others generate speech with the face that is saying it. The lips are not a second product. The room tone is not a bed you will duck later. The model decided what the mouth was doing in the same denoise as the words. That is why native audio is the default whenever we can see the speaker.

    TTS is a speech-model feature. You write the line, you pick or clone a voice, you lock the style, you get a file. The picture does not know about it. If you then lay that file on a silent generate, the mouth will not match unless you add a lipsync pass — a third product, with its own credit line. If you lay it on a native-audio generate, you now have two voices fighting one face.

    The catalog page that lists who can even do the first job is best model with native audio. The catalog page for the second job is AI text-to-speech. Open both before you generate. Do not discover at picture lock that the model you picked was silent by default.

    The decision

    Ask two questions, in this order.

    1. Is the mouth the shot?

    If yes — talking head, presenter, UGC face, anyone whose lips we see — native audio is the default. Generating silent and stacking TTS is how you get a face that is chewing while the voice is already on the next clause. A lipsync repair can close that gap. It is a repair. It costs more than picking an audio-capable model the first time.

    If no — product turntable, landscape, hands, b-roll, a talking shot we never cut to — native in-shot speech is wasted. TTS, a music bed, or silence plus captions is the cleaner file.

    2. Does this voice have to be the same person next week?

    If yes — a founder clone, a series narrator, a brand voice that legal signed off — TTS (or a cloned voice into a talking-avatar pipeline) is the lock. Native audio will give you a plausible speaker every time, and a different plausible speaker every time. That is fine for a one-off. It is a recast for episode 12.

    If no — a one-shot explainer, a character who only exists in this clip, a gag where the voice is the world — let the video model talk.

    Brief Mouth in frame? Same voice next week? Route
    Talking UGC ad, one SKU, one take Yes No Native audio. Write the line into the prompt.
    Founder series, 30 episodes Yes Yes Clone the founder. Avatar or lipsync. Do not let Veo invent a cousin.
    Product hero, no face No Irrelevant Silent or TTS bed. Native dialogue is noise.
    Faceless explainer, off-screen VO No Yes TTS with a style lock. Captions on the picture.
    In-world character, this clip only Yes No Native. The room and the mouth are one event.

    Do not mix both on one mouth

    The failure that looks professional in the timeline and amateur on the phone:

    You generate a Veo take with native dialogue because the model is good at talking. You do not like the voice. You duck the native track and drop a TTS line on top. The lips still belong to the first voice. The new voice is late, or a different vowel, or a different person. Viewers may not name the artefact. They will feel it.

    Fixes that are not fixes:

    • Regenerating native until the voice "sounds like" the founder. You will spend the series chasing a timbre the model does not own.
    • Generating silent "so we can add TTS later" on a talking shot. You threw away the only free lipsync you were going to get.
    • Native on the A-roll, TTS on the same person's B-roll. That is two people.

    If you must replace native speech, replace the picture of the mouth as well — a silent generate, a talking-avatar pass, or a lipsync dub — so one source owns the face. If you must keep the native take, keep the native voice. Taste is not a reason to stack.

    What to write, once you have picked

    Native: put the line in the prompt, in quotes, with the room. Name ambience if it matters. Leaving sound unspecified is how you get a saxophone in a skincare ad.

    TTS: lock the voice and the style before the batch. The series problem is not one bad line. It is 30 lines that each sound like a slightly different morning. Style-locking is the series tool; native audio is the in-shot tool.

    Captions are not a third audio path. They are how the file works muted. Burn them on after the voice is decided, not instead of deciding.

    FAQ

    Can I generate native audio and then swap the voice in the edit?

    Not on the same mouth. Swapping the track leaves the original lips. If the voice is the deliverable, pick TTS (or a clone) before you generate the face. If the take is the deliverable, keep the native voice.

    When is silent-plus-TTS actually the right talking shot?

    When the voice is a contractual object — a named founder, a licensed talent, a series bible — and the video model cannot take that voice as an input. Then you generate (or film) for the face and drive the mouth from the audio. That is a lipsync job, not a "we'll fix it in captions" job.

    Does native audio mean I should never use TTS on a video with a person in it?

    No. Off-screen narration over a person who is not speaking is TTS's home turf: a founder voice over hands, a VO over b-roll, a recap over a silent reaction. The rule is one source per mouth, not one source per video.

    Which models should I even consider for native dialogue?

    Start at best model with native audio and then filter for the shot: always-on talking heads still lean Veo; multilingual talking shots lean Happy Horse; long single takes lean Seedance 2.5. Elo is not the filter. Whether the model emits speech in the same pass as the mouth is.