Guides

    Put the camera first: prompt order that holds framing

    Leading a video prompt with shot size and angle changes the framing you get back. A five-slot shot-card skeleton you can fill for any shot, in order.

    Versely Team9 min read

    Take a prompt that mostly works, move the words "wide shot, low angle" from the last line to the first, change nothing else, and generate it again. The framing changes. Not the subject, not the lighting, not the grade — the framing. That is the whole argument for prompt order, and it is cheap enough to verify that you should not take my word for it.

    The skeleton worth defaulting to is [Cinematography] + [Subject] + [Action] + [Context] + [Style & ambiance]: cinematography in the first slot, before anyone has said who is in the shot. That is not an arbitrary arrangement. It matches how a shot list is written on set, where size and angle get named before the actor, because framing is the container everything else has to fit inside.

    The five slots

    Here is the skeleton, with the thing each slot is actually for and the thing people keep putting in it by mistake.

    Slot Belongs here Does not belong here
    1. Cinematography Shot size, angle, camera movement or explicit stillness, focal length Mood adjectives, grade words
    2. Subject Who or what, with two or three identifying details Backstory, personality, motivation
    3. Action One verb, one beat A sequence of events with a beginning and an end
    4. Context Where, when, and the named light source Style words, camera words
    5. Style and ambiance Grade, film stock, texture, era A second camera instruction

    The discipline is that each slot gets its own clause and nothing leaks across. Camera language belongs in slot one and nowhere else. If "handheld" shows up in slot five as a texture word after you asked for a locked frame in slot one, you have written a contradiction, and a contradiction gets resolved differently on every take.

    Why the first slot wins

    The practical read is that a model fills in a camera whether or not you specify one. Left unspecified, it picks a default that is usually a mid-shot with some drift. When your camera instruction arrives at the end of the prompt, it is not setting the framing — it is trying to override a framing the model has already committed to while it was busy rendering everything you said before it. Sometimes that override lands. Often it does not, and the clip comes back looking like your camera line was never there.

    Put the same words first and they are no longer an override. They are the premise everything else gets built on top of. This is the same failure surface described in camera instructions models silently ignore: the instruction parses fine and still does not execute. Position is one of the levers that changes whether it executes.

    Shot size is the highest-yield single token you can move to the front. "Extreme wide," "medium close-up," "over-the-shoulder," "two-shot" — these change the picture more reliably than any other camera word, and they cost you four syllables. Angle is second: "low angle," "high angle," "bird's eye," "dutch angle." Movement is third and least reliable, because camera control is a real exposed parameter on only part of the catalog and prose everywhere else.

    Filling the slots without padding them

    A prompt that has all the right words in the wrong order:

    A woman in a red coat walking through a rainy Tokyo side street at night, neon reflections everywhere, cinematic, moody, atmospheric, shot on 35mm, slow dolly in

    Everything the model needs is in there. The camera arrives last, behind five words that mean nothing on their own. Rebuilt as a shot card:

    Medium shot, slight low angle, 35mm, slow dolly in. A woman in a red coat, hood down, hair wet. She walks toward camera without breaking stride. A narrow Tokyo side street at night, lit by pink neon signage on the left wall and reflected off wet asphalt. Muted cyan and magenta grade, fine film grain.

    Same scene, roughly the same word count. What changed: camera first, one action verb instead of a gerund that never resolves, a named light source with a direction instead of "neon reflections everywhere," and the three mood adjectives deleted. "Cinematic," "moody," and "atmospheric" mostly bias the grade toward desaturation and contrast. They do not put a light anywhere. One concrete lighting descriptor beats five style adjectives, every time.

    The action slot rewards restraint hardest. One verb, one beat. "She walks toward camera" is a shot. "She walks toward camera, stops, looks back, then keeps going" is four shots crammed into one generation, and the model will compress or drop most of it. If you need the sequence, that is a first-and-last-frame problem or a multi-clip problem, not a prompt-length problem.

    The optional sixth slot

    On models that generate audio in the same pass rather than as a layer, the shot card grows a slot at the end. The convention that holds up is one audio element per sentence: dialogue inside quotation marks (A woman says, "We have to leave now."), effects labelled (SFX: thunder cracks in the distance), and background stated on its own (Ambient noise: rain on corrugated metal). The full treatment, including where native audio beats a separate voice pass, is in dialogue and audio prompts.

    The restraint rule here is two named foley events, maximum. Specific and countable beats atmospheric: "espresso machine hiss" and "glass-on-marble click" will land; "busy cafe sounds" will produce an undifferentiated bed. Let the model fill the ambient layer on its own and spend your specificity on the two events the edit actually needs to hear.

    When the skeleton should shrink

    The five-slot card is a default, not a law. Some model families take a fully populated scene-level brief without choking; others do better with a short declarative shot line plus a reference image, and degrade when handed a long one. Density is genuinely model-specific in a way that order is not, and it is worth testing on its own — short prompts or long prompts works through the length side of the question by model class.

    What does not change between families is where the camera line sits. When you need to shrink the card, drop slots from the back — style first, then context — rather than demoting cinematography to make room. A prompt with three slots and the camera in front holds framing better than a prompt with five slots and the camera in the tail.

    Testing your own order in one sitting

    1. Write one shot card properly, all five slots, camera first.
    2. Write the same content as a single run-on sentence with the camera language at the end. Same words, different arrangement.
    3. Run both three times each on the same model. Three takes, not one — a single clip tells you what happened once, and order effects show up as a difference in hit rate, not as a difference in any individual take.
    4. Score one thing only: did you get the shot size and angle you named? Not whether it looks good. Ignore the grade, ignore the subject, ignore whether the neon is pretty.
    5. Repeat on a second model before you generalize. The agent can fan one prompt across several named models in a single request, which makes the second half of this test roughly as much work as the first half.

    The result is a hit rate you can act on. If camera-first lands the named framing five times out of six and camera-last lands it twice, that is the end of the debate for that model, and you have a skeleton you can fill for every shot after it. The rest of the vocabulary — shot-size ladder, lens language, placement — is covered in shot composition prompts, and the movement primitives themselves in camera movement prompts for AI video. This post is only about what order to say them in.

    FAQ

    Does prompt order matter as much for images as for video?

    Less, but it still matters. A still image has no camera move to schedule, so the highest-value front-loaded token is shot size and angle rather than movement. The bigger win in image prompts is the same restraint rule: one concrete framing instruction ahead of the subject beats a pile of style adjectives behind it.

    Should I put camera language at the front even on models with a camera-control parameter?

    Yes, and check the parameter too. Where a model exposes real camera control as a setting, the setting is authoritative and your prose should agree with it rather than fight it. Where it does not, the prose is all you have, and position is one of the few levers you can pull. Versely's model catalog lists which controls each video model actually exposes.

    What if my shot genuinely needs two camera moves?

    Then it is two shots. Compound camera instructions are where prompt adherence degrades fastest — one component usually wins and the other quietly disappears. Generate them as separate clips and cut, or describe the move as a transition between two endpoints instead of as an adjective.

    Do I need all five slots on every prompt?

    No. Missing slots get filled by the model's defaults, which is fine when you do not care about that dimension. The slot you should never leave empty is the first one, because that default is the one you are most likely to notice and least likely to be happy with.