Guides

    Consistency Collapse: Diagnosing Drift Between Shots

    'It doesn't look consistent' isn't one problem. It's four — identity, wardrobe, set and lighting — each with a different cause and a different fix.

    Versely Team8 min read

    "These shots don't match" is a diagnosis that feels complete and isn't. Lay out a batch of AI-generated shots meant to be the same production and the mismatch is almost never one thing — it's some combination of four distinct failures that happen to look like a single vague problem from a distance. Naming them separately matters, because each one has a different cause and a different fix, and treating all four as one undifferentiated "consistency issue" tends to mean fixing the one you noticed first and shipping the other three broken.

    Why "inconsistent" is really four different failures

    Each generation in a batch is, structurally, an independent event unless something explicitly ties it to the others. What varies between shots is whatever the prompt left underspecified — and different categories of detail get left underspecified for different reasons, which is why they drift differently and need different fixes.

    Identity drift is a character or subject whose face, proportions, or defining features shift shot to shot — the viewer registers a different person playing the same role. This is the most visually disqualifying failure because human perception is unusually tuned to faces specifically; even a subtle jaw or eye-shape shift reads instantly as "someone else," in a way a subtle color shift elsewhere in the frame doesn't.

    Wardrobe drift is the same character with an outfit that quietly changes — a red shirt reads as maroon two shots later, a jacket picks up buttons it didn't have. Clothing has far more descriptive degrees of freedom than a face does — exact hue, texture, fold pattern, fit — and a prompt that nails the character's face rarely pins every one of those details with equal precision, so wardrobe drifts even when identity survives.

    Set drift is the environment itself changing shape between shots meant to share a location — furniture that moves, a room that gains or loses geometry, an object that isn't quite the object it was in the previous shot. This one is documented in the wild in a way the others often aren't: an Adobe Firefly community bug report describes a chair regenerated across repeated attempts sprouting "spider-like legs, different chair bases, wheeled platforms," and switching material from fabric to metal — the same described object, reinterpreted differently nearly every time, because nothing in a text description pins down geometry the way a reference image would.

    Lighting drift is the subtlest and the most commonly missed: same character, same wardrobe, same set, but the mood, color temperature, or direction of the light shifts between shots that are supposed to share a continuous scene. It's the least-specified property of all, because most prompts describe a subject, sometimes a wardrobe, occasionally a set — and rarely explicit light direction or color temperature. What goes unspecified gets re-decided every generation, and lighting is usually what goes most unspecified.

    The diagnostic order that actually saves time

    Reviewing a batch for all four failures at once is how you miss the ones that aren't visually loud. A fixed order catches more, faster:

    1. Identity first. It's the most disqualifying failure on its own and the fastest to spot — if the face or subject has drifted, nothing else about the shot matters yet, because the whole shot needs regenerating regardless of how good the wardrobe or lighting turned out.
    2. Wardrobe second. Check it right after identity because both usually share the same fix: a subject reference image. If you're already regenerating for identity, checking wardrobe in the same pass catches a second failure with one review.
    3. Set third. This one takes closer inspection to catch — a chair with an extra leg or a background object that's subtly different geometry doesn't jump out the way a wrong face does, which is exactly why it needs a deliberate check rather than a glance, per the pattern the Firefly report describes.
    4. Lighting last, but don't skip it. It's the hardest to catch because any single shot's lighting usually looks fine in isolation — the problem only appears in sequence, when a scene that's supposed to be continuous jumps from cool to warm between cuts. Review the batch back to back, not shot by shot, to catch this one.

    That order isn't arbitrary — it goes from "most disqualifying and easiest to spot" to "least disqualifying alone and hardest to spot," which means a rushed review that only gets through the first two or three items still catches the failures most likely to sink the whole piece.

    Different failures, different fixes

    Our own character consistency guidance already gets at part of this: "keep wardrobe, lighting and lens language identical across shots — those read as identity too," which is exactly why identity and wardrobe drift are worth checking together — they usually share a root cause (nothing pinned the subject down beyond words) and share a fix (supplying pixels, not more adjectives, via a reference image). Set drift needs the same category of fix aimed at a different target: a reference image of the location or object itself, not just the character, since a text description of a chair is exactly as under-specified as a text description of a face. Lighting drift is the one failure a reference image doesn't automatically solve — it needs explicit, repeated lighting language in every prompt in the sequence, the same discipline a spec sentence brings to any other under-specified property.

    Worth a clean distinction here too: none of this is temporal consistency, which is a different failure entirely — flicker and instability within a single continuous clip, frame to frame, rather than drift between separate generations. A shot can be perfectly stable internally and still be the wrong shot to cut next to the one before it. Consistency collapse is a between-shots problem; temporal consistency is a within-shot one, and they need checking separately.

    Reference-to-video models: multiple anchors in one generation

    The practical lever behind three of the four fixes above is the same: give the model pixels instead of description. Reference-to-video models take this further than a single reference image — several in Versely's catalog accept multiple simultaneous references per generation, which matters because it means a character reference and a set reference can both steer the same generation rather than competing for the one reference slot a simpler model gives you. Minimax H3 Reference to Video and Seedance 2.0 Reference to Video both accept up to 9 reference images (plus reference video and audio); Wan 2.7 Reference to Video accepts up to 5 reference images and 5 reference videos; Kling O3 Pro Reference to Video and VEO 3.1 Reference to Video accept several reference images each. Practically, that headroom is what lets you anchor identity, wardrobe, and set in the same request instead of picking just one to fix and hoping the others hold.

    Building the fix in Versely

    1. Build canonical reference assets before generating a batch, not after noticing drift. Setting up reusable characters and products opens upload cards for a character, product, prop, or location — and for a full cast, several at once rather than one at a time — so identity, wardrobe, and set each have a dedicated reference asset to anchor to before the first shot is even generated.
    2. Reuse those assets by name across every shot in the sequence, rather than re-describing the subject, outfit, or location fresh in each prompt — the saved asset is reusable by key across scenes, which is the whole point of setting it up once.
    3. Write the lighting language explicitly into every prompt in the sequence, since it's the one failure a reference asset doesn't automatically cover — repeat the same time-of-day, color-temperature, and direction language shot to shot rather than assuming it carries over.
    4. Review in the diagnostic order above, batch, not shot by shot — identity and wardrobe first since they share a fix, set third with closer inspection, lighting last across the full sequence in one pass.
    5. Shortlist a reference-to-video model by current rank on Versely's reference-to-video comparison before committing a long sequence to one — capability varies enough between models that it's worth checking before, not after, a batch.

    FAQ

    Is consistency collapse the same problem as temporal consistency (flicker)?

    No — they're adjacent but distinct. Consistency collapse is drift between separate shots or generations (a character, outfit, set, or lighting that changes from one clip to the next). Temporal consistency is instability within one continuous clip, frame to frame. A clip can be internally stable and still not match the shot before it.

    Which of the four failures is worth fixing first if I only have time for one pass?

    Identity, without much competition — it's both the most visually disqualifying failure on its own and usually shares a fix with wardrobe drift, so addressing it with a character reference asset tends to improve wardrobe consistency in the same pass.

    Do I need a different reference image for the character, the wardrobe, and the set, or can one image cover everything?

    Ideally separate references for each, especially on models that accept multiple reference inputs per generation — a single image asked to anchor a face, an outfit, and a location simultaneously is being asked to do the job of three, and set or wardrobe details tend to lose precedence to whatever the model weights as the primary subject.