Veo burned in subtitles you never asked for
Quoted text and contractions trigger Veo auto-subtitles. The prompt rewrite that stops it, and why writing 'no subtitles' makes it worse.
You asked for a woman setting a cup down. You got the cup, the room, the light, and a row of garbled captions along the bottom of the frame that you never wrote. Writing "no subtitles" in the next prompt makes it worse. That is not you failing to be emphatic. It is how Veo reads a prompt.
The behaviour was documented on Veo 3 within weeks of the May 2025 launch. MIT Technology Review reported it in July 2025: clips with dialogue often arrived with nonsensical burned-in subtitles even when the prompt asked for none, and Google Labs' Discord support told users the subtitles could be triggered by speech. Josh Woodward of Google Labs posted on 9 June 2025 that the team had developed fixes to reduce the gibberish text; users were still logging the issue a month later. The current catalog model on Versely is Veo 3.1. The same prompt-reading pattern is the one you still write around: quoted text and contractions get treated as spoken dialogue, and naming "subtitles" in order to forbid them puts the concept in the conditioning.
This is a prompt rewrite, not a setting. There is no "captions off" control on the generation. If you want real captions, add them afterwards as a timeline layer. If you want a clean frame, stop giving Veo the cues that mean "this is a subtitled talking-head clip."
Why the captions appear, and why "no subtitles" summons them
Two mechanisms, stacked.
Training data is full of burned-in social captions. Veo was built to produce video with native audio and speech. A large share of human talking-head video on the public web has captions composited into the frame, not sitting in a sidecar file. Shuo Niu, quoted in the Technology Review piece, put the mechanism plainly: if the model is rewarded for resembling human-created video, and that video includes subtitles, the model learns that subtitles are part of the resemblance.
Negation is weakly modelled. Telling a generative model not to do something names the thing. Tuhin Chakrabarty, also quoted in that reporting, is the general statement; the Veo case is the specific one. "No subtitles," "no captions," "don't add text" all put those nouns into the prompt. The word "no" is not an operator. It is more tokens. That is the reason what to write instead of a negative prompt exists as a standing rule.
A third, Veo-specific trigger sits on top: the prompt looks like a script. Quotation marks are how you mark spoken lines. Contractions (I'm, don't, it's) are how spoken English is written. A quoted, attributed line is the documented way to request dialogue on Veo 3.1. It is also a reliable way to request the picture of speech, which on this model includes the captions that usually come with it.
So the failure is not mysterious. You wrote a talking-head prompt in the orthography of a talking-head video. The model completed the pattern, including the lower third.
The rewrite that actually stops it
Convert every negative into a description of the frame and the soundtrack. Remove the script cues if you do not want speech. If you do want speech, keep the line and drop the punctuation that marks it as a caption.
1. Delete "no subtitles," "no captions," "no text on screen." Replace with a positive description of a clean picture.
The frame is a clean photograph with an empty lower third. No lettering is drawn on the image. Spoken audio only, if any.
Even that last sentence is safer as a description of the audio than as a ban on type. Prefer: "The words are heard, not drawn."
2. Remove quotation marks unless you are deliberately requesting a spoken line. A phrase in quotes is a line of dialogue to this model. If the quotes were for emphasis or for a product name, rewrite them out.
| You wrote | Rewrite |
|---|---|
| A barista says "Last one of the day" | A barista speaking, last one of the day, you got lucky |
| The sign reads "OPEN" | A shop window with the word OPEN painted on the glass |
| She whispers "I'm here" | She whispers that she is here |
The middle row is the easy miss. Quoted on-screen type (a sign, a label, a phone screen) gets read as dialogue, which then pulls captions for a line nobody spoke. If the shot needs a real word in the scene, you are in on-screen text territory, and the honest move is a clean plate plus a type layer, not a quoted string inside a Veo prompt.
3. Expand contractions. Write I am, do not, it is, we are. The spoken line can still be informal in the performance direction ("casual, offhand") without being spelled as a caption.
4. Specify the soundtrack as a set of sounds, not as an absence. "Silent scene, ambient room tone, refrigerator hum, distant traffic." Give the audio channel something to be, so it does not reach for speech-plus-captions as the default fill. If the mouth should not move, say so in the picture: "lips closed, still face, small blinks only."
5. Do not put the negative field to work as a sentence. Where Veo exposes a negative field, list nouns, not instructions: subtitles, captions, text, lower third, not "no subtitles." The words "no" and "don't" are noise in that channel. If you have already cleaned the positive prompt, a short noun list is optional insurance, not the fix.
A before and after for the common case, a clip you wanted without speech:
Weak (summons captions and talking):
Close-up of a woman at a kitchen counter, no talking, no subtitles, no captions. She says nothing. "Quiet morning energy."
Strong:
Close-up of a woman at a kitchen counter, lips closed, eyes on the kettle. Static shot, tripod. Soft morning window light. Audio: kettle beginning to boil, distant birds, room tone. The frame is a clean photograph with an empty lower third.
A before and after for the case where you do want the line, and you do not want it drawn:
Weak (quoted line, contraction, and a ban):
Medium shot of a barista sliding a cup across the counter, saying: "I'm out after this." No subtitles.
Strong:
Medium shot of a barista sliding a cup across the counter, speaking in a low voice that she is out after this. Warm morning light, gentle cafe chatter. The words are heard on the soundtrack only. The picture has an empty lower third.
If the line has to be word-perfect, generate the picture without speech and put a real voice on afterwards. Off-script paraphrase is a documented Veo 3.1 behaviour of its own; a quoted string is not a contract. The Veo 3.1 prompting guide is the right reference when you want native dialogue. Use it, then apply this rewrite so the native dialogue does not arrive with a burned-in transcript. Add "empty lower third, words heard not drawn" as a standing clause on any talking shot.
If the captions already landed
Do not re-roll the whole clip as the first move. You would be re-rolling the parts that already work.
Add real captions on purpose, on a clean plate, if you need them. Burned-in captions that you control are a timeline job. Add captions to a video transcribes the actual soundtrack and composites styled captions you picked. If the model already burned a nonsense track into the frame, this tool cannot sit on top of it cleanly; you need a clean plate first.
Crop only as a last resort, and only if the captions sit in a band you can afford to lose. On a 9:16 talking head a lower-third crop often takes the chin or the product with it. Try the rewrite and a short retake first.
Retake the clip with the rewritten prompt. Keep it to a few seconds. If the first retake still lettered the frame, look at whether a quoted product name or a contraction survived the rewrite.
What not to do: stack "NO SUBTITLES NO TEXT NO CAPTIONS" in caps. Emphasis is not negation. You have now mentioned the concept six times.
FAQ
I need the character to speak. How do I keep the words off the picture?
Write the line without quotation marks, expand the contractions, and describe the picture as a clean frame with an empty lower third. Put the performance in the audio direction ("low voice, close to the microphone") rather than in a quoted script. If the wording has to be exact, generate a silent picture and add a real voice and real captions on the timeline, where you control the type.
Is this only Veo 3, or does 3.1 do it too?
The Technology Review reporting is about Veo 3 at launch in 2025. Versely's catalog model is Veo 3.1. Write as if 3.1 still has the same prior: quotes and contractions as speech cues, burned-in social captions in the training mix, and negation that names the thing you wanted gone. You will be right more often than if you assume a silent fix landed.
Why does "no subtitles" make it worse?
Because those two words are not a filter. They are a mention. The model conditions on subtitles. The "no" does not reliably invert the mention, and on a model whose training set treats captions as part of talking-head video, a mention is enough. Describe the clean frame instead.
Can I just crop the captions off and ship the clip?
Sometimes, if they sit in a letterbox you were going to lose anyway. On a full-bleed 9:16 shot they usually sit on the speaker's chest or the product, and a crop to remove them is a new composition. Cropping is also how you discover the captions were not only at the bottom. Retake with the rewrite. It is the cheaper way to keep the shot you already liked.