Guides

    Keep the prompt structure, vary only the action

    Multi-shot consistency comes from freezing the prompt skeleton and changing one clause. The template, the fields that must never move, and how people break it.

    Versely Team9 min read

    Write six shots of the same character as six freshly composed prompts and you will get six slightly different productions. Write them as one prompt with one clause swapped and you get a sequence. The difference is not prompt quality. It is that a re-written prompt silently re-decides everything you didn't deliberately restate, and a frozen prompt has nothing left to re-decide except the thing you changed.

    This is the cheapest consistency mechanism available, it needs no extra tooling, and it stacks with references rather than competing with them.

    The skeleton, and why cinematography goes first

    The slot order worth defaulting to:

    [Cinematography] + [Subject] + [Action] + [Context] + [Style & ambiance]
    

    Leading with cinematography rather than subject runs against instinct — most people open with who is in the shot. But shot size is among the highest-yield tokens in a video prompt. Put it first and it governs composition; bury it at the end of a long prompt and it is one of the first things to get diluted.

    Once you accept that order, the multi-shot technique falls out of it. Five slots, four frozen, one variable. A character holds recognisably across a run of shots when the prompt structure stays identical between them and only the action changes.

    Which fields move and which never do:

    Slot In a sequence Why
    Cinematography Frozen per scene Shot size and lens drive framing; changing them mid-scene reads as a different production, not a different shot
    Subject Frozen, word for word Every re-description is a re-roll of appearance; even a synonym shifts it
    Action The variable This is the only clause you are allowed to rewrite
    Context Frozen The set is one set; re-describing it lets geometry drift
    Style & ambiance Frozen, word for word Lighting and grade language is the most commonly under-specified slot, so it is the first thing to move if you let it

    "Frozen" means literally identical text, not equivalent text. "A woman in a navy jacket" and "a woman wearing a navy jacket" are the same sentence to you and two different conditioning inputs to the model. Copy and paste; do not retype.

    If you want a change of shot size — and you should, a scene of six identical framings is dull — change it deliberately at a planned point and treat that as a new frozen block. Shots one through three are the medium-shot block, four through six are the close-up block, and inside each block the cinematography text does not move.

    The template

    [CINEMATOGRAPHY] Medium shot, 35mm, shallow depth of field, camera
    locked off and entirely motionless.
    
    [SUBJECT] A woman in her thirties, faded navy canvas jacket over a
    plain white t-shirt, hair tied back, steel watch on the left wrist.
    
    [ACTION] <<< THE ONLY LINE THAT CHANGES >>>
    
    [CONTEXT] A narrow galley kitchen, late afternoon, window light
    entering from camera left, subject at roughly 45 degrees to the source.
    
    [STYLE] Naturalistic, muted palette, soft shadow from the large window
    source, no colour grade beyond a gentle lift in the shadows.
    

    Four shots built on it — note that everything outside the action line is byte-identical across all four:

    Shot Action clause
    1 She sets a white ceramic mug down on the counter and lets go of the handle.
    2 She turns her head toward the window and holds there.
    3 She picks the mug back up with her right hand and brings it to chest height.
    4 She takes one step out of frame to camera right, leaving the counter empty.

    Each action is a single physical event with a start and an end. Multi-event clauses — "she picks up the mug, drinks, sets it down and walks out" — ask a short generation to schedule four beats, and the model compresses or drops some. One beat per shot; the shot list carries the sequence.

    Note what the cinematography line says explicitly: camera locked off and entirely motionless. Models carry a strong prior toward adding movement, and omitting camera language does not give you a static frame — it gives you whatever drift the model felt like.

    Four ways people break the discipline without noticing

    Adding one adjective to the subject slot. You look at shot four, decide she should read a little more tired, and add "looking tired" to the subject line. That is a re-described subject, and the face comes back different. Emotional state is an action, not an attribute — put it in the action clause where it belongs, or set it with a reference.

    Reordering under pressure. Halfway through a sequence someone moves the style line above the action line because it "reads better." The slots are positional cues as much as content, and the order that produced shots one to three is part of what made them match.

    Writing contradictory camera language. static handheld shot and locked-off tracking shot are internally inconsistent, and the model resolves them by effectively picking one at random — which means the same prompt gives you different answers on different runs. Contradictions are the fastest way to lose repeatability.

    Reaching for a negative prompt. The old reflex is to append "no blur, no distortion, no extra limbs" when a shot comes back wrong. On models that run without classifier-free guidance the negative field may do nothing at all. Even where negative prompts still function, instructive phrasing tends to leak the negated noun back in. The replacement is positive re-specification: describe the frame you want — an empty counter, a bare wall, hands resting flat — rather than listing what should not be in it.

    Where prose runs out

    Freezing the skeleton buys you a great deal and then stops. Two limits are worth knowing before you blame your writing.

    Speed and easing are not inferable from text. "Slow dolly in" has no defined rate, so the model re-estimates it every generation and your three shots move at three speeds. When a move has to be repeatable across a sequence, parameterised camera control beats prose, because it locks the move as configuration rather than as an adjective. The motion-control shortlist is a reasonable starting point.

    A single generation can only schedule so much. If a shot needs several beats, the tool is not a longer action clause. Some models respond to timestamped prompting — beats written as explicit [00:00-00:02] ranges inside one generation — which is worth a test on whatever you are using rather than an assumption. Where a shot's end state is narratively load-bearing, first-last-frame generation lets you specify both ends and hand the model only the motion between them. The chaining version of that is in keyframe chaining for longer scenes.

    Model behaviour differs here in ways worth testing rather than assuming. Some models take a long, scene-level structured brief happily; others do better with a short declarative shot line carried by reference images. Which is which moves with releases, so run your own skeleton against a couple of candidates rather than inheriting anyone's ranking, including one written today.

    Running a frozen skeleton in practice

    Keep the skeleton in a text file, not in your head or in a chat window. The whole method depends on the frozen text being genuinely identical, and the most common source of drift is a human retyping a line they could have pasted.

    From there the mechanics are ordinary. The AI video generator takes the assembled prompt per shot, and the agent chat can fan a single prompt across several named models in one request — the fastest way to find out whether your skeleton survives on a model you have not used before. Generations are billed in credits; the ladder is on pricing.

    Structure discipline is one mechanism among several and not the strongest one; references bind harder. Four consistency mechanisms and when each fails compares them. The reason to start here is that freezing a skeleton takes a minute and makes every other mechanism easier to evaluate, because you have removed the variable of your own prose.

    FAQ

    If I have a strong character reference, do I still need to freeze the prompt?

    Yes, for everything the reference does not cover. A reference binds appearance inside its crop. The set, the lighting, the lens and the grade are still decided by text on every run, and those drift just as visibly as a face does across a cut.

    How long can the action clause be?

    One physical event, usually one sentence. If you need a second sentence you probably need a second shot. Short generations schedule few events well and many events badly, and splitting gives you a usable cut rather than a compressed beat.

    Can I change the context line if the character walks into another room?

    That is a new scene, so yes — but write it as a deliberate new frozen block rather than an edit to the old one, and keep the subject and style lines byte-identical across the boundary. Changing one slot at a scene break is normal; changing one mid-scene is drift.

    What do I do when a shot comes back wrong but the prompt is right?

    Re-roll before you rewrite. A frozen skeleton makes this diagnosable: if the same text produces a good shot on the second attempt, you hit a bad sample. If three identical runs fail the same way, the prompt is genuinely underspecified — and you now know which single clause to change, because everything else was held constant.