Guides

    Dialogue and Audio Prompts for Native-Audio Models

    Dialogue and audio prompts for native-audio AI video models: scripting spoken lines, cueing sound effects and ambience, and when to use TTS instead.

    Versely Team7 min read

    The newest generation of video models doesn't just render the scene — it performs it. Prompt VEO 3.1 or Sora 2 with a line of dialogue and you get the voice, the lip movement, the room tone, and the footsteps, all synthesized together in one pass. It's the biggest workflow change in AI video since image-to-video, and it comes with a prompting discipline nobody teaches: your prompt is now a script, with spoken lines, sound cues, and stage directions — and the models are opinionated about how those should be written. Here's the format that works, plus the honest decision about when native audio is the wrong tool.

    Recording dialogue and sound

    Script format: quote the line, describe the voice

    Native-audio models treat quoted text as speech to perform, the same way typography models treat quoted text as copy to render. The reliable pattern has three parts — speaker, delivery, line:

    A woman in a yellow apron looks into the camera and says warmly: "Three ingredients. Ten minutes. Let's go."

    Rules that raise the hit rate:

    • Quote dialogue exactly. Unquoted speech ("she introduces the recipe") makes the model improvise lines — sometimes usable, never on-script. If the words matter, quote them.
    • Keep lines short. One or two sentences per clip is the reliability zone for a 5–10 second generation. Long monologues get rushed, truncated, or mumbled through. Break speeches across clips and cut.
    • Direct the delivery, briefly. One adverb or short phrase before the line: "says warmly," "whispers urgently," "announces with mock seriousness." Delivery direction is honored surprisingly well and is the difference between a read and a performance.
    • Describe the voice if it matters. "A deep, gravelly voice," "a bright, fast-talking presenter." You won't get a specific voice actor — you're steering register and energy, and consistency across clips is loose. That looseness matters later; hold that thought.

    For two-person exchanges: label turns explicitly ("The man asks: '…' The woman replies: '…'"), keep it to one exchange per clip, and make the speakers visually distinct so lips land on the right face.

    Cueing sound effects and ambience

    Dialogue is half the soundtrack. The other half — effects and atmosphere — responds to its own prompt vocabulary:

    • Ambience beds: name the environment's sound explicitly. "Busy café ambience, low chatter and espresso machine hiss." Models generate convincing room tone when told what the room sounds like — and near-silence when not.
    • Spot effects tied to action: attach the sound to the visible event. "She sets the mug down with a soft ceramic clink." Effects synced to on-screen actions are the native-audio party trick; free-floating effect requests ("add cool sounds") produce mud.
    • Sound as story beat: "The room is silent except for a ticking clock" is an audio composition instruction, and strong native-audio models honor the exception structure — silence plus one feature — beautifully. It's the audio equivalent of negative space.
    • What to leave out: music. Prompted background music is where native audio is weakest — generic, poorly mixed, and legally ambiguous territory at output. Generate clips with dialogue and effects only, then add a proper track in the edit where you control the mix.

    Which models can do what here varies more than any spec sheet suggests — native-audio video models explained covers the landscape, and the working shortlist lives in the best AI models for native audio. On the product side, Vidu Q3's audio is distinctive enough that it carries entire product stories without a voiceover session.

    Native audio vs TTS + lipsync: the real decision

    Native audio is magic in a demo and a tradeoff in a pipeline. The honest fork:

    Factor Native audio TTS + lipsync
    Speed to first result One generation, done Two or three steps
    Voice consistency across clips Loose — voice drifts between generations Exact — same voice, every clip, forever
    Script revisions Regenerate whole clip (video changes too) Regenerate audio only; video untouched
    Long scripts Truncation risk past ~2 sentences Unlimited
    Sound synced to action Excellent, automatic Manual in the edit
    Brand voice / voice cloning Not controllable Fully controllable

    The pattern that falls out: native audio for moments, TTS for narrators. A one-off character line, a reaction with room tone, an ambient product scene — native, one pass. A recurring host, a 60-second explainer, anything a client will revise line-by-line — generate the voice with text-to-speech (or a cloned brand voice), then sync it to footage with Sync Lipsync 2.0. Revision economics decide it: when a stakeholder changes one word of a native-audio clip, you re-roll the entire video and hope the visuals hold; with the TTS path you re-render one sentence of audio.

    Hybrids are legitimate and common: native audio for the scene's ambience and incidental lines, TTS narrator laid over the top. Just duck the native bed under the narrator in the mix.

    Failure modes worth knowing in advance

    • Wrong words spoken. Shorten the line and simplify vocabulary; proper nouns and brand names are mispronunciation magnets. If the brand name must be said, the TTS path pronounces it the same way every time — native audio re-rolls it as a lottery.
    • The dreaded caption echo. Some models render your quoted dialogue as on-screen subtitles unprompted. Append "no on-screen text or subtitles" — you'll add proper styled captions in post, where you control them.
    • Voice changes mid-series. Not fixable with prompts; native voices aren't stable identities across generations. Series work belongs on the TTS path, full stop.
    • Muddy overlapping audio. One speaker, one ambience, one or two spot effects per clip. Audio complexity budgets are real, and they're smaller than visual ones.
    • Silent output from an audio-capable model. Usually an under-specified prompt — no quoted line, no named ambience. Native audio needs audio instructions, not just permission.

    FAQ

    How do I make a model actually speak my exact line?

    Quote it, keep it under two sentences, and attach it to a described speaker: "she says warmly: 'line here.'" Unquoted paraphrases get improvised. For lines that must survive revision cycles or exceed a couple of sentences, switch to TTS plus lipsync instead of fighting truncation.

    Can I choose or keep a specific voice with native audio?

    You can steer register — deep, bright, gravelly, fast-talking — but not pin an identity, and the voice will drift between generations. For a consistent recurring voice (a host, a brand narrator, a cloned voice), generate speech with a TTS system and lipsync it to the footage.

    Why is background music from native-audio models bad advice?

    It's the weakest output of the audio stack — generically composed and baked into the mix, so you can't adjust levels against dialogue later. Prompt for dialogue, ambience, and spot effects only, then add a licensed or generated music track in the edit where you control ducking and levels.

    How much audio can one clip carry?

    One speaker, one ambience bed, and one or two action-synced effects is the reliable ceiling for a short generation. Stacked speakers or dense soundscapes turn to mud. For busy scenes, generate the audio in layers — native ambience in the clip, narration and music in post.

    Which is cheaper overall: native audio or the TTS pipeline?

    For one-off clips, native usually wins — one generation instead of three steps. For anything revised or serialized, TTS wins decisively: script changes re-render seconds of audio instead of whole video clips, and voice consistency comes free. Price the revision loop, not the first take.

    Write your next clip as a three-line script — speaker, delivery, quoted line, plus one ambience cue — and run it on a native-audio model in Versely. If you find yourself revising it twice, you already know which pipeline it belongs in.