The model added dialogue you never wrote
Native-audio models treat 'no talking' as a cue to talk. Convert every negative into a positive audio state, with worked before-and-after prompts.
You wrote "no talking, no dialogue, silent scene" and the person on screen still opened their mouth. That is not a stubborn model. Naming talking is how you put talking into the conditioning. The word "no" is not an operator that subtracts it.
Native-audio video models generate picture and soundtrack in one pass. If the prompt does not pin what we should hear, the model fills the bed from the prior that most of its training clips share: someone speaking, often with type burned into the frame. The fix is mechanical. Stop listing what the soundtrack must not contain. Describe the audio state you actually want.
"No" is a token, not a switch
A negative prompt is supposed to be a second conditioning that pushes away from a concept. A negation sitting inside the main prompt is not that. It is more tokens in the same embedding. "No talking" still contains "talking". The model has a strong visual and sonic prior for talking, and a weak, poorly trained prior for the linguistic trick of inversion.
That is why the failure is so consistent. "No subtitles" produces garbled captions. "She doesn't speak" produces a line. "Mute" produces muttering. You activated the concept and then hoped a tiny grammatical particle would cancel it. It does not.
The same mechanism shows up in stills ("no text" yielding letter-shaped noise). On video it is louder because native audio has to decide what the clip sounds like as well as what it looks like. A silent result is not the default. Silence is a specific request, and it has to be made in the positive: closed mouth, named room tone, no spoken turn.
Native audio fills whatever you leave blank
Until recently the pipeline was: generate a silent clip, then source or generate sound and align it. A native-audio model does both together, which is why footsteps can land on the footfall. It also means an underspecified soundtrack is not empty. It is improvised.
Leave audio unmentioned and the model chooses. The training prior for short, photoreal clips is social video: a face, a line, often burned-in captions. That is a documented behaviour on Google's Veo 3. MIT Technology Review reported in July 2025 that Veo 3 kept slapping garbled, nonsensical captions onto clips even when users asked for none, and that negative instructions were the wrong tool. Tuhin Chakrabarty, quoted in that piece, put the general rule plainly: negative prompts are usually less effective than positive ones.
So the prompt has to own the soundtrack the way it owns the shot. Three slots, filled every time:
- Mouth state. Closed, listening, chewing, breathing through the nose. If the mouth is visible and you do not specify it, speech is the cheap guess.
- Ambience. Name the sources: refrigerator hum, kettle, rail clack, HVAC, cloth against a worktop. Mood words ("quiet", "tense") are not sound.
- Who takes a turn. If nobody speaks, say the subject is listening. If someone does speak, quote the line. Unquoted "she explains the recipe" is an invitation to improvise.
The existing dialogue and audio prompting guide covers the case where you want a performed line. This page is the inverse: you want the picture, and you want the room, and you do not want a performance.
Convert every negative into an audio state
Rewrite the exclusion as a description of the bed. The test is simple. If you deleted every "no" and "don't" from the sentence, would the remaining nouns still point at speech, music, or type? If yes, those nouns are doing the damage.
| You wrote | What arrives | Write instead |
|---|---|---|
| no talking, no dialogue | A performed line | Mouth closed, listening. Audio: refrigerator hum, kettle beginning to steam |
| she doesn't speak | She speaks anyway | She watches the window, jaw relaxed, a slow blink. No spoken turn; room tone only |
| no music / no score | A generic bed of music | Production sound only: cloth rustle, distant traffic, a clock tick |
| no subtitles, no captions | Garbled burn-in | Clean picture; the soundtrack lives in the audio; type added in the edit |
| no crowd, empty street | Crowd murmur on the track | Dawn street, one occupant, wind across shopfronts, a distant bird |
| mute / silent | Muttering, or a score | Locked-off, closed mouth, ambient bed of HVAC and fridge compressor |
| don't say anything | A line about not saying anything | He holds the pause. Breath only. The kettle is the loudest thing in the room |
Worked pair, kitchen scene.
Before:
Medium shot of a woman in a sunlit kitchen, preparing coffee, no talking, no music, no subtitles, cinematic.
After:
Medium shot of a woman in a sunlit kitchen, standing at the counter, mouth closed, eyes on the kettle. Slow blink. Static camera, 50mm. Audio: refrigerator hum, water coming to a boil, ceramic cup set down. Clean picture; captions added later in the edit.
The after version never names talking, music, or subtitles. It names a mouth state, three sound sources, and where type will be handled (not in the generation). That last clause matters. If you need captions, generate without them and burn a real transcript in post, where you control spelling and placement.
Worked pair, train carriage.
Before:
Close-up of a man on a commuter train, he doesn't speak, no dialogue, quiet.
After:
Close-up of a man on a commuter train, jaw relaxed, eyes on the window. Audio: rail clack, carriage rattle, an unintelligible distant announcement. He is listening. Mouth closed throughout.
"Quiet" was a mood. "Rail clack, carriage rattle" is a bed. The model can render a bed. It cannot reliably invert a noun.
The Veo special case: quotes, contractions, type
VEO 3.1 is the current native-audio Veo in the catalog. The subtitle habit was documented on Veo 3 in 2025. Three extra rules sit on top of the rewrite above.
Do not put quotation marks in the prompt unless you want those words spoken. Quoted text is treated as a line to perform, the same way a typography model treats quoted text as copy to render. A sign, a label, or a bit of first-person scene description in quotes will get read aloud. If the words are not a line, do not quote them.
Expand contractions if the words must appear but must not be performed. Practitioner reports against Veo are consistent on this: "I'm here" is more likely to be spoken than "I am here". The contraction is dialogue-shaped. Write the full forms, or take the first-person sentence out of the prompt entirely and put it on a prop you will composite later.
Do not fight burned-in type with "no subtitles". That phrase names subtitles. Describe a clean frame and move captioning to post.
If the shot needs a controlled voice (a brand line, a host, anything you will revise), do not ask the video model to perform it. Generate the picture with a closed mouth and a named ambience, then add a dedicated read with text-to-speech plus lipsync. Native audio is the wrong tool for a line you must own.
A practical sequence on the video generator:
- Write the shot with no audio words at all. Generate once. Listen.
- Whatever arrived that you did not want (a line, a score, a caption, crowd), name the replacement bed, not the offender.
- Specify mouth state if a face is in frame.
- Strip every remaining "no" / "don't" / "without" that still sits next to a speech, music, or type noun.
- Only if a quoted line is required, add it in quotes, short, attributed. That is a different prompt, and it should not share a sentence with "no talking".
The VEO 3.1 prompting guide is the right companion when the model is Veo and you are filling the audio slot on purpose. The rule on this page applies across native-audio models: the soundtrack is part of the brief, and inversion is not how you write it.
FAQ
Why does a prompt that says "silent" still produce muttering?
"Silent" is a mood, not a bed. The model still has to emit an audio track, and the cheap fill is a voice-like texture. Name the actual sources you will accept (HVAC, fridge, cloth, distant traffic) and the mouth state. Silence in a native-audio clip is usually "one quiet source", not digital zero.
Should I put "dialogue, talking, speech" in the negative prompt field?
Not as the first move, and not while the positive prompt still contains those words. On models with a working negative channel, a short noun list can discourage an object class. On native-audio video it is unreliable, and on some families the field is inert. Rewrite the positive prompt first. If the same intrusion appears on two takes after that, then a bare noun in the negative field is a test, not a workflow.
How do I get a talking head without the model inventing the line?
Quote the line, keep it to one or two sentences, and attach it to a described speaker: she says, evenly: "The oven is at 180." Unquoted paraphrases get improvised. If the words have to survive a revision cycle, generate the picture silent and perform the line in TTS plus lipsync so a one-word change does not regenerate the face.
Will switching models stop the extra speech?
Sometimes the burn-in rate drops. The underlying habit does not, because any native-audio model trained on social video has seen talking faces with type in the frame. A cleaner prompt travels. A model swap with the same "no talking" clause usually does not.