Guides

    Four shots that make a scene read as a scene

    Establishing, medium, insert and reverse turn a pile of clips into a place. The coverage set, and how to write four prompts that agree with each other.

    Versely Team9 min read

    Generate eight clips of a coffee shop and you have eight coffee shops. Each one is plausible. None of them is the same room, and cutting between them produces something the audience will not read as a location — they will read it as a montage, or as a mistake, depending on how much you were asking of it.

    A scene is not a quantity of footage. It is a set of shots that agree about where they are, when they are, and which way things face. Four is the smallest set that gets you there: an establishing, a medium, an insert, and a reverse. More is often nicer, but four is the floor, and the specific four matter.

    What each of the four does

    Shot Framing The job it does The thing it must not change
    Establishing Wide, environment dominant Tells the audience where they are and how big it is Architecture, light source, time of day
    Medium Waist-up or two-shot Carries the action and most of the dialogue Wardrobe, subject position relative to the light
    Insert Close-up or macro on an object Gives the scene a specific detail to be about Surface, material, colour temperature
    Reverse Opposite angle on the same action Proves the room has two directions Which side the light comes from — it must flip

    The establishing shot is undervalued because on its own it looks like stock footage. Its real job is not to be watched — it is to be the reference the other three are built from.

    The insert is the cheapest and most reliable shot in the set. It has the smallest amount of frame to get wrong, and it is what you reach for when two other shots refuse to match. Generate more inserts than you think you need.

    The reverse is the one that breaks, and the one that does the most work: a location only ever photographed from one direction is a backdrop, not a place.

    Write one scene block and paste it into all four

    The single most effective technique here is boring: write the room once, then reuse that exact text in all four prompts and change only the shot line and the action. Models hold a location or a character noticeably better when the surrounding text stays fixed than when each shot is written fresh. Rewriting the description of the room four different ways is the most common way people accidentally generate four different rooms.

    So build a block, once, and treat it as constant text:

    Scene block (constant across all four shots):
    Small independent coffee shop, mid-morning. Exposed brick on the far
    wall, dark timber counter, brass fittings. Large window on the east
    side. Daylight enters from camera left in the master, subject at 45
    degrees to the source. Warm practical pendant lamps over the counter.
    Muted palette, warm neutrals. Shot on a 35mm lens.
    

    Then each shot is that block plus a shot line and an action line. Lead with the cinematography — putting the shot size first reliably changes framing, and putting it last often gets it dropped altogether. Most published video prompt guides converge on the same skeleton: camera, then subject, then action, then context, then look.

    Shot 1 — Establishing:
    Wide shot, camera entirely motionless, eye level. A barista wipes down
    the counter at the far end of the room. [scene block]
    
    Shot 2 — Medium:
    Medium shot, slight low angle, camera entirely motionless. The barista
    sets a cup on the counter and looks up. [scene block]
    
    Shot 3 — Insert:
    Extreme close-up, macro, shallow depth of field. Steam rising from a
    white ceramic cup on dark timber. [scene block]
    
    Shot 4 — Reverse:
    Medium shot from behind the counter looking back toward the window,
    camera entirely motionless. Daylight now enters from camera RIGHT.
    The brick wall is out of frame; the window fills the background.
    [scene block, with the light direction line replaced]
    

    The reverse changes exactly two things: the light comes from the opposite side of frame, and the background is whatever was behind the camera in the master. Get those right and it reads as the same room. Get them wrong — light still from the left, brick wall still behind the subject — and you have a second coffee shop that resembles the first, which is worse than an obvious mismatch because the geometry feels subtly broken.

    The order to generate in

    Generation order is not cut order. Cut order is establishing, medium, insert, reverse. Generation order is:

    1. Establishing first. Run it until you get one you would happily put on screen. This shot sets the look for everything else, and time spent here is repaid three times over.
    2. Harvest a frame from it. Pull a still and use it as a reference image for the remaining three. Reference-first beats re-prompting: a locked image constrains far more of the picture than any amount of adjective matching, and it is the same mechanism that keeps a face stable across a shot list. The broader version of this discipline is covered in shot coverage sets.
    3. Insert next, because it is fast and it tells you whether your palette and colour temperature carried over. If the insert comes back with a different white balance, fix the block before you spend anything on the reverse.
    4. Medium.
    5. Reverse last, with the two flips applied, and expect to run it more times than the other three combined.

    If the subject is a person who appears in more than one of the four, lock their identity in a single portrait first and feed that alongside the scene reference. Character consistency is a separate problem from location consistency and it does not solve itself just because the room matched.

    Say "the camera does not move"

    Video models have a strong prior toward adding movement. Omitting camera language does not give you a static shot — it gives you an unrequested drift, and four shots each drifting a different amount is a scene that feels seasick.

    For a four-shot set, make at least the establishing and the reverse locked off. Write it explicitly: the camera is entirely motionless. If you want a move, name the mechanism rather than the effect — "the camera physically travels toward the subject" instead of "closer," which models resolve as either a dolly or a zoom at random. And accept that prose cannot specify speed or easing: "slow push" is re-estimated every generation, so if you need the identical move on two shots, a parameter-based camera control setting is the only way to get it twice.

    A useful alternative for the one moving shot in the set: describe the move as a transition between two endpoints using first and last frame rather than as an adjective. "Smooth arc from front-facing to behind the counter" is much more reliably executed when both ends of the arc are supplied as images.

    Assembling it

    Four shots at three to five seconds each is twelve to twenty seconds of usable scene — more than most social edits need, and enough to carry a beat inside a longer piece. Keep the clips short; that length is also where models stay temporally stable, so the format and the technical limit agree.

    Cut establishing → medium → insert → reverse for a scene that opens out, or insert → medium → establishing → reverse for one that reveals. Assemble in the video editor, or push the set straight through merge if all four already share a frame rate and resolution. If they don't, fix that first — a mismatched set reads as bad continuity when it is actually a format problem.

    The block scales the obvious way: a three-location piece is three blocks of four. What you should not do is generate twelve shots against twelve freshly written descriptions and hope the edit sorts it out.

    FAQ

    Why does the reverse fail more than the other three?

    Because it asks the model to infer something it was never shown: what the room looks like from a position that has not appeared in any frame. The other three photograph the same half of the space; the reverse photographs the other half. Give it help — name the background explicitly, flip the light direction, and if the space has a distinctive feature behind camera, put it in the block from the start rather than inventing it at reverse time.

    Can I get all four out of one generation?

    Sometimes. Veo 3.1 supports timestamp prompting, where you segment a single generation into blocks like [00:00-00:02] and [00:02-00:04] and direct a short sequence of shots inside one clip. It keeps the location consistent by construction. The trade-off is control: you cannot regenerate shot three without regenerating all four, so it suits drafts and social cuts better than work that goes through revisions.

    How many inserts should I actually generate?

    Two or three per scene, of different objects. Inserts are the shortest, most forgiving shots in the set and they are what you cut to when a match fails somewhere else. An insert between a medium and a reverse gives the audience a moment where they are not comparing the two, which buys you a surprising amount of tolerance on the reverse.

    Does this work for a location that doesn't exist yet?

    It works better. A real location has a ground truth you can be caught contradicting; an invented one only has to agree with itself, and the scene block is the artefact that makes it agree — as long as you write the block before the first shot rather than reverse-engineering it from whatever the first generation produced.