Guides

    VEO 3.1 Prompting: The Complete Guide

    VEO 3.1 prompting guide: a shot-based prompt structure, dialogue and native audio cues, camera language, and before/after fixes for weak prompts.

    Versely Team8 min read

    The same VEO 3.1 credits can buy you a muddy, aimless clip or a shot that looks deliberately directed — and the difference is almost never the model having a good day. It's whether your prompt reads like a wish or like a shot description. VEO 3.1 rewards prompts written the way a director briefs a cinematographer: one subject, one action, one camera move, one lighting condition, and — because VEO generates native audio and dialogue — one clear instruction about what we should hear. This guide gives you a repeatable structure for VEO 3.1 prompting, worked examples you can adapt, and the failure modes that waste the most credits.

    Cinematic robot portrait under studio lighting

    The anatomy of a VEO 3.1 prompt

    Treat every prompt as six slots, in roughly this order:

    1. Shot type — "close-up," "wide establishing shot," "over-the-shoulder."
    2. Subject — one clearly described subject with 2–3 concrete attributes.
    3. Action — a single continuous action that fits a short clip.
    4. Setting — place, time of day, weather.
    5. Camera + lens feel — movement and style: "slow dolly-in," "handheld," "35mm feel."
    6. Light + audio — lighting condition, then sound: ambience, effects, or dialogue.

    A filled example:

    Close-up of a weathered fisherman in a yellow raincoat, coiling rope on the deck of a trawler. Cold overcast morning, North Atlantic. Slow dolly-in, shallow depth of field. Gray diffuse light, sea spray on the lens. Audio: wind, creaking hull, distant gulls.

    Notice what's absent: no "beautiful," "cinematic masterpiece," "8K award-winning." VEO responds to concrete nouns and physical description far better than quality adjectives, which mostly add noise.

    Dialogue and native audio: prompt what you hear

    Native audio is VEO 3.1's signature advantage, and most weak prompts simply ignore it — leaving the soundscape to chance. Give audio its own sentence, and be literal.

    Ambience: "Audio: rain on a tin roof, low thunder, no music." The "no music" matters — if you don't want a score, say so.

    Dialogue: put the line in quotes and attribute it:

    Medium shot of a barista in a sunlit café, sliding a cup across the counter, saying: "Last one of the day — you got lucky." Warm morning light, gentle café chatter in the background.

    Keep spoken lines short — one or two sentences a clip. Long monologues drift out of sync with performance and eat your shot's runtime. For multi-line scenes, generate one line per clip and cut them together; it's more controllable and each retake is cheaper.

    Common audio failure: prompting mood ("tense atmosphere") instead of sound sources. The model needs things that make sound — "a clock ticking, fluorescent hum" beats "tense atmosphere" every time.

    Camera language VEO understands

    VEO 3.1 handles standard cinematography vocabulary well, but only one move per clip. Stacking "dolly-in while orbiting then crane up" usually produces a wobbly compromise. Pick from a small reliable set:

    You want Say Avoid saying
    Slow push toward subject "slow dolly-in" "zoom in dramatically"
    Circling the subject "orbit shot, 90 degrees around" "camera spins around fast"
    Documentary energy "handheld, slight shake" "shaky cam"
    Reveal of scale "crane up from street level to rooftops" "epic reveal"
    Locked, composed frame "static shot, tripod" (omitting camera entirely)

    That last row is underrated: if you say nothing about the camera, VEO invents a move, and it's often the wrong one. "Static shot" is a real instruction, not a default.

    Reference-to-video: locking identity across shots

    For brand and character work, VEO 3.1 reference-to-video accepts reference images so the subject stays consistent across generated shots. Prompting changes in one key way: stop re-describing what the reference already shows. The reference carries identity; your prompt should spend its words on what's new — action, setting, camera, audio.

    Before: "A woman with shoulder-length brown hair, green eyes, freckles, wearing the blue jacket from the photo, walks through a market…" (fighting the reference with text).

    After: "She walks through a crowded night market, tasting food from a stall. Tracking shot alongside her. Lantern light, sizzling wok audio." (the reference does the identity work).

    This is the backbone of multi-shot consistency: one reference set, several prompts that vary only scene and action. For the model's broader capabilities beyond prompting, see the complete VEO 3.1 guide.

    Before/after: three weak prompts, fixed

    Weak: "A cool cyberpunk city at night, amazing detail, cinematic." Fixed: "Wide establishing shot of a rain-soaked neon street at night, crowds with umbrellas, steam rising from vents. Slow push-in at street level. Reflective wet asphalt, magenta and cyan signage. Audio: rain, crowd murmur, distant sirens." Why it works: the weak version has no subject, no action, no camera, no audio — four empty slots the model fills randomly.

    Weak: "A dog runs and then jumps in a lake and then shakes off water and looks at camera." Fixed: "A golden retriever sprints down a wooden dock and leaps into a lake, water exploding on impact. Slow-motion feel, tracking shot from the side. Late afternoon sun. Audio: paws on wood, a big splash." Why it works: one clip, one beat. Chained actions ("and then… and then…") are the single most common VEO failure — the model compresses or drops beats. One action per generation, sequence in the edit.

    Weak: "Product video for a perfume bottle, luxury style." Fixed: "Macro shot of an amber glass perfume bottle on black marble, a single droplet sliding down the glass. Slow orbit. One hard rim light from the left, everything else falls to black. Audio: soft room tone, no music." Why it works: "luxury" is an outcome; hard rim light on black marble is an instruction.

    Failure modes and fast fixes

    • Morphing or extra limbs mid-action → the action is too complex for the clip; simplify to one motion.
    • Ignored dialogue → line wasn't in quotes, or competed with heavy action; give speech its own calmer shot.
    • Unwanted music → you never specified audio; add "no music" plus explicit sound sources.
    • Style drift across a series → your style words vary between prompts; keep an identical style sentence in every prompt of the set, or move identity into references.
    • Text and signage garbled → keep on-screen text minimal or add it in post with overlays instead of asking the model to render it.

    When a prompt works, save it as a template with the subject and setting slots blanked. A personal library of five proven structures outperforms writing fresh prose every time — and running the same structure across other models on /models tells you quickly where VEO 3.1 is genuinely the best pick for the shot.

    FAQ

    What makes VEO 3.1 prompting different from other video models?

    Native audio and dialogue. Most models generate silent footage; VEO builds a soundscape, so prompts that specify audio sources and quoted dialogue use a whole dimension other prompts waste. Structure-wise it also rewards director-style shot descriptions over adjective piles.

    How do I get VEO 3.1 to generate dialogue correctly?

    Put the exact line in quotation marks, attribute it to a described speaker, and keep it to one or two short sentences per clip. Prompt the delivery context too — "saying quietly," "shouting over the wind" — and generate longer conversations one line at a time.

    Why do my VEO 3.1 clips look generic?

    Almost always empty slots: no camera instruction, no lighting condition, no audio. The model fills unspecified slots with averages, and averages look generic. Filling all six slots — shot, subject, action, setting, camera, light+audio — is the fastest quality jump available.

    Should I use text-to-video or reference-to-video in VEO 3.1?

    Text-to-video for one-off shots where identity doesn't need to persist. Reference-to-video whenever a character, product, or style must stay consistent across multiple shots — and once you're using references, spend prompt words on action and setting rather than re-describing the subject.

    Can I test the same prompt on VEO 3.1 and other models?

    Yes — that's the practical advantage of a multi-model platform. Run one structured prompt through VEO 3.1, Kling O3, and Seedance 2.0 in Versely, compare results side by side, and check the live ELO rankings to see which model currently leads for your shot type.

    Ready to put the structure to work? Open Versely's AI video generator, pick VEO 3.1, and run the six-slot template on your next shot — then run the identical prompt on a second model and let the results argue it out.