VEO 3.1 Prompting: The Complete Guide
VEO 3.1 shot prompts, native audio cues, camera language, and fixes. Run the structure on Versely.
On Versely, the same VEO 3.1 credits can buy you a muddy, aimless clip or a shot that looks deliberately directed, and the difference is almost never the model having a good day. It's whether your prompt reads like a wish or like a shot description. VEO 3.1 rewards prompts written the way a director briefs a cinematographer: one subject, one action, one camera move, one lighting condition, and, because VEO generates native audio and dialogue, one clear instruction about what we should hear. This guide gives you a repeatable structure for VEO 3.1 prompting, worked examples you can adapt, and the failure modes that waste the most credits.
The anatomy of a VEO 3.1 prompt
Treat every prompt as six slots, in roughly this order:
- Shot type — "close-up," "wide establishing shot," "over-the-shoulder."
- Subject — one clearly described subject with 2–3 concrete attributes.
- Action — a single continuous action that fits a short clip.
- Setting — place, time of day, weather.
- Camera + lens feel — movement and style: "slow dolly-in," "handheld," "35mm feel."
- Light + audio — lighting condition, then sound: ambience, effects, or dialogue.
A filled example:
Close-up of a weathered fisherman in a yellow raincoat, coiling rope on the deck of a trawler. Cold overcast morning, North Atlantic. Slow dolly-in, shallow depth of field. Gray diffuse light, sea spray on the lens. Audio: wind, creaking hull, distant gulls.
Notice what's absent: no "beautiful," "cinematic masterpiece," "8K award-winning." VEO responds to concrete nouns and physical description far better than quality adjectives, which mostly add noise.
Dialogue and native audio: prompt what you hear
Native audio is VEO 3.1's signature advantage, and most weak prompts simply ignore it — leaving the soundscape to chance. Give audio its own sentence, and be literal.
Ambience: "Audio: rain on a tin roof, low thunder, no music." The "no music" matters — if you don't want a score, say so.
Dialogue: put the line in quotes and attribute it:
Medium shot of a barista in a sunlit café, sliding a cup across the counter, saying: "Last one of the day — you got lucky." Warm morning light, gentle café chatter in the background.
Keep spoken lines short — one or two sentences a clip. Long monologues drift out of sync with performance and eat your shot's runtime. For multi-line scenes, generate one line per clip and cut them together; it's more controllable and each retake is cheaper.
Common audio failure: prompting mood ("tense atmosphere") instead of sound sources. The model needs things that make sound — "a clock ticking, fluorescent hum" beats "tense atmosphere" every time.
Camera language VEO understands
VEO 3.1 handles standard cinematography vocabulary well, but only one move per clip. Stacking "dolly-in while orbiting then crane up" usually produces a wobbly compromise. Pick from a small reliable set:
| You want | Say | Avoid saying |
|---|---|---|
| Slow push toward subject | "slow dolly-in" | "zoom in dramatically" |
| Circling the subject | "orbit shot, 90 degrees around" | "camera spins around fast" |
| Documentary energy | "handheld, slight shake" | "shaky cam" |
| Reveal of scale | "crane up from street level to rooftops" | "epic reveal" |
| Locked, composed frame | "static shot, tripod" | (omitting camera entirely) |
That last row is underrated: if you say nothing about the camera, VEO invents a move, and it's often the wrong one. "Static shot" is a real instruction, not a default.
Reference-to-video: locking identity across shots
For brand and character work, VEO 3.1 reference-to-video accepts reference images so the subject stays consistent across generated shots. Prompting changes in one key way: stop re-describing what the reference already shows. The reference carries identity; your prompt should spend its words on what's new — action, setting, camera, audio.
Before: "A woman with shoulder-length brown hair, green eyes, freckles, wearing the blue jacket from the photo, walks through a market…" (fighting the reference with text).
After: "She walks through a crowded night market, tasting food from a stall. Tracking shot alongside her. Lantern light, sizzling wok audio." (the reference does the identity work).
This is the backbone of multi-shot consistency: one reference set, several prompts that vary only scene and action. For the model's broader capabilities beyond prompting, see the complete VEO 3.1 guide.
Before/after: three weak prompts, fixed
Weak: "A cool cyberpunk city at night, amazing detail, cinematic." Fixed: "Wide establishing shot of a rain-soaked neon street at night, crowds with umbrellas, steam rising from vents. Slow push-in at street level. Reflective wet asphalt, magenta and cyan signage. Audio: rain, crowd murmur, distant sirens." Why it works: the weak version has no subject, no action, no camera, no audio — four empty slots the model fills randomly.
Weak: "A dog runs and then jumps in a lake and then shakes off water and looks at camera." Fixed: "A golden retriever sprints down a wooden dock and leaps into a lake, water exploding on impact. Slow-motion feel, tracking shot from the side. Late afternoon sun. Audio: paws on wood, a big splash." Why it works: one clip, one beat. Chained actions ("and then… and then…") are the single most common VEO failure — the model compresses or drops beats. One action per generation, sequence in the edit.
Weak: "Product video for a perfume bottle, luxury style." Fixed: "Macro shot of an amber glass perfume bottle on black marble, a single droplet sliding down the glass. Slow orbit. One hard rim light from the left, everything else falls to black. Audio: soft room tone, no music." Why it works: "luxury" is an outcome; hard rim light on black marble is an instruction.
Failure modes and fast fixes
- Morphing or extra limbs mid-action → the action is too complex for the clip; simplify to one motion.
- Ignored dialogue → line wasn't in quotes, or competed with heavy action; give speech its own calmer shot.
- Unwanted music → you never specified audio; add "no music" plus explicit sound sources.
- Style drift across a series → your style words vary between prompts; keep an identical style sentence in every prompt of the set, or move identity into references.
- Text and signage garbled → keep on-screen text minimal or add it in post with overlays instead of asking the model to render it.
When a prompt works, save it as a template with the subject and setting slots blanked. A personal library of five proven structures outperforms writing fresh prose every time — and running the same structure across other models on the model catalog tells you quickly where VEO 3.1 is genuinely the best pick for the shot.
Skip the dialect: brief the Versely agent
You can learn the six-slot structure. On Versely you can also skip memorizing every VEO dialect and brief the agent instead of writing prompts: content type, photos or references, the edits you will accept, target platforms, and a budget ceiling. Name the row when it matters ("use VEO 3.1 reference-to-video for product identity"). The agent plans the job; you approve the plan instead of babysitting slot order.
A prompt still wins when you are hand-tuning one hero take with locked audio and camera. A brief wins when the job is bigger than one generate: variants, captions, and a post path with a spend cap.
After the take: edit, post, collections
When the VEO pass is close enough:
- Light trim and pacing in the studio editor if a beat overruns. Do not ask the model to be CapCut.
- Burn mute-proof captions with /tools/ai-caption-generator or the on-device /free-tools/burn-captions path when the take is silent-friendly, or keep native dialogue and caption only what helps mute viewers.
- Upload or schedule to the social accounts Versely already connects for your workspace. If a network is not connected, export the mp4 and post from the native app.
- Save keepers into a Versely collection so the next brief reuses the same stills, refs, and winning prompt lines.
Honest limit: VEO 3.1 is not a full multi-scene movie editor. Board separate generates when you need separate setups, then stitch. One action and one camera move per clip is still the product.
FAQ
What makes VEO 3.1 prompting different from other video models?
Native audio and dialogue. Most models generate silent footage; VEO builds a soundscape, so prompts that specify audio sources and quoted dialogue use a whole dimension other prompts waste. Structure-wise it also rewards director-style shot descriptions over adjective piles.
How do I get VEO 3.1 to generate dialogue correctly?
Put the exact line in quotation marks, attribute it to a described speaker, and keep it to one or two short sentences per clip. Prompt the delivery context too — "saying quietly," "shouting over the wind" — and generate longer conversations one line at a time.
Why do my VEO 3.1 clips look generic?
Almost always empty slots: no camera instruction, no lighting condition, no audio. The model fills unspecified slots with averages, and averages look generic. Filling all six slots — shot, subject, action, setting, camera, light+audio — is the fastest quality jump available.
Should I use text-to-video or reference-to-video in VEO 3.1?
Text-to-video for one-off shots where identity doesn't need to persist. Reference-to-video whenever a character, product, or style must stay consistent across multiple shots — and once you're using references, spend prompt words on action and setting rather than re-describing the subject.
Can I test the same prompt on VEO 3.1 and other models?
Yes — that's the practical advantage of a multi-model platform. Run one structured prompt through VEO 3.1, Kling O3, and Seedance 2.0 in Versely, compare results side by side, and check the live ELO rankings to see which model currently leads for your shot type.
Ready to put the structure to work? Open Versely's AI video generator, pick VEO 3.1, and run the six-slot template on your next shot — then run the identical prompt on a second model and let the results argue it out.