Guides

    Dialogue coverage: two-shots and reverse angles

    Two people talking needs four shots, not one. The coverage set, the eyeline rule that makes reverses cut together, and the reference discipline behind both.

    Versely Team9 min read

    Generate "two people arguing across a kitchen table" and you get a competent clip that cannot be cut into anything. It is one angle, both faces held at the same size for the whole duration, and the moment you need to shorten a line or land a reaction you discover there is nowhere to go. The scene isn't badly generated. It is under-covered.

    Dialogue is the one case where a single angle is never enough: a conversation is an exchange, and an exchange has to be seen from both sides. Live-action crews solve it with a coverage pass — wide, then each performer's side, then a detail. Generated dialogue needs that same pass plus one discipline live action gets for free, because when two performers are generated independently nothing guarantees they are looking at each other.

    The four shots

    Four generations cover almost any two-hander. Everything beyond this is refinement.

    Shot What's in frame What it's for
    Two-shot Both performers, waist-up or wider Establishes who is where, and the geography you're stuck with afterwards
    Over-the-shoulder (A on B) B's face, A's shoulder in the foreground B's lines, with A's presence still in frame
    Reverse (B on A) A's face, B's shoulder in the foreground A's lines and reactions
    Insert Hands, an object, a screen — no faces Pace change, and a place to cut when the two faces don't quite match

    The two-shot is the shot to generate first, not because it's the best shot but because it decides everything else. Whichever side of the table it puts each performer on is the arrangement the other three shots have to honour. Generate the close angles first and you will find yourself building a wide that has to agree with two shots that already disagree with each other.

    The insert earns its place twice over. It contains no face, so it has no identity to hold and rarely needs a reroll — and it hands the edit an escape hatch. Wherever the eyelines are half a degree off or a wardrobe detail changed between takes, a two-second cutaway to a hand on a mug makes the problem disappear.

    Eyelines are the thing that actually breaks

    This is the part that separates dialogue coverage from a nine-angle coverage set built around one subject. With one subject, every angle is valid on its own. With two, each shot has to be consistent with a spatial relationship the model was never told about.

    Set it up once and repeat it verbatim. Say A is on the left and B is on the right, facing each other:

    • In every shot, A occupies the left half of frame and looks toward screen right.
    • In every shot, B occupies the right half of frame and looks toward screen left.
    • The camera stays on one side of the pair for all four shots.

    Eyelines match by being opposite, not identical. If both performers look screen-left in their respective singles, they read as staring at a third thing off-camera and the reverse cut fails — the audience won't articulate why, they'll just feel the scene has gone slack. Get this pair of directions right and two clips generated hours apart will cut together as if they were shot on the same afternoon.

    Height matters too, and it's easier to forget. If A is seated and B is standing, A's eyeline goes up and B's goes down, in every single shot. State it: she is seated, looking up at him; he is standing, looking down at her. A model will happily render both performers at a neutral horizontal eyeline and you will only notice at assembly.

    One tempting shortcut you should not take: do not generate the over-the-shoulder once and horizontally flip it in post to produce the reverse. It flips hair partings, wardrobe asymmetry, watches, and any text in the background. It also gives you the same performance twice, which is exactly what a reverse angle is supposed to avoid.

    Lock both faces before you generate anything

    Prompt-only descriptions of two people will not survive four separate generations. "A woman in her forties with dark curly hair" describes a category, not a person, and each clip samples a different member of that category. The 2026 working method is reference-first: build the identity once as an image, then feed that image into every shot.

    Practically that means:

    1. Generate or supply one clean portrait per performer — even lighting, torso-up crop, neutral expression, high resolution. Mixed lighting and inconsistent crops across your two references are what cause the model to average two faces into a third.
    2. Build the outfit and framing variants you need from those portraits rather than re-describing the person each time.
    3. Route the shots through reference-to-video rather than plain text-to-video. Seventeen of the catalog's active video models take that path, including Veo 3.1's reference-to-video mode.

    The over-the-shoulder shots are the awkward ones, because they need two people in frame — one facing camera, one from behind. The shoulder in the foreground is soft and partial, so it is far more forgiving than a face, but it still has to carry the right hair, the right jacket, and the right shoulder height. Feed both references where the model accepts more than one; where it accepts only one, give it the performer whose face is on screen and describe the foreground shoulder in text.

    Keep the prompt structure identical between shots

    Once the geography and the references are set, the highest-yield habit is boring: write one prompt skeleton and change only the shot line.

    [SHOT] Medium over-the-shoulder, camera behind the man's right shoulder,
           favouring the woman.
    [SUBJECTS] The woman is seated at frame right, looking toward screen left.
               The man's shoulder fills the left foreground, out of focus.
    [SETTING] Small kitchen, late afternoon, window light from camera left.
    [ACTION] She sets down her mug and says, "You said you'd tell me first."
    

    For the reverse, the [SETTING] block is copied character for character. Only [SHOT], the screen positions, and the line change. This isn't superstition. A fresh description of the same kitchen is a different set of tokens, and a different set of tokens samples a different kitchen; holding the wording fixed removes that variable entirely. Rewriting the setting in new words between shots is how cabinets change halfway through a scene.

    Voices, and why the reverse sounds like a different person

    Native-audio models generate the voice with the frames, in one pass. That is excellent for a single clip and a problem for a four-shot scene, because four independent generations produce four independent voice castings. The performer's face can hold perfectly and the reverse still sounds like it was dubbed by someone else.

    Two ways out. Either keep all of one character's dialogue inside a single generation and cut around it — which is why scene direction sized to clip length is worth reading before you break a scene into shots — or separate the audio entirely: generate the picture silent, produce the dialogue once with a fixed voice, and attach it with lipsync. The second path costs an extra step per shot and is the only one that guarantees the same voice across four angles. The trade-offs between the two are laid out in more detail in the dialogue and audio prompting guide.

    Assembling and checking it

    Cut the four shots together in the AI video editor before you commit to a finished render. The timeline is EDL-based, so it stays re-renderable rather than being a one-shot export, and a preview: true pass returns 480p at no credit cost with a short per-user cooldown between requests. The final export is charged once regardless of how many clips are on the timeline — the exact shape of that is in the previews and final export breakdown.

    480p is plenty for the checks that matter here: eyeline direction survives any downscale, and so does the mismatch that makes a reverse cut fail. Watch it once through at full speed with sound on. A scene that reads as two people in a room reads that way at 480p, and one that doesn't won't be rescued by resolution.

    FAQ

    How many generations does a two-person dialogue scene really need?

    Four is the working minimum: a two-shot, an over-the-shoulder each way, and one insert. That gets you a scene you can cut, tighten, and re-pace. Five if either performer has a reaction the script actually depends on, since a reaction usually wants a tighter single than the over-the-shoulder gives you. Beyond that you are refining rather than covering.

    Why do my two characters look like they aren't talking to each other?

    Almost always eyeline direction. Each clip is sampled independently, so nothing carries the spatial relationship between them, and if both singles happen to look the same way on screen the pair stops reading as an exchange. Fix it by stating screen position and look direction explicitly in every prompt — one performer on the left looking right, the other on the right looking left — rather than describing them as "facing each other", which a model can satisfy without agreeing with the other shot.

    Can I flip a shot horizontally to make the reverse?

    Not on anything with a face and wardrobe in it. Flipping reverses hair partings, buttons, jewellery, and any legible text in the background, and it hands you the same performance twice when a reverse exists to give you a different one. Generate the reverse properly; use the flip only on abstract inserts where nothing in frame is asymmetric.

    Should the dialogue be generated with the video or added afterwards?

    For a single clip, native audio in one pass is simpler and better integrated. For a multi-shot scene, generating the picture silent and attaching one consistently voiced track afterwards is the safer default, because four independent generations will cast four subtly different voices for the same character.