Guides

    Reference hygiene: why your refs fight each other

    Mixed lighting, crops and resolutions make a model average two looks into a third person. The hard spec for a reference set that holds one identity.

    Versely Team9 min read

    You upload six photos of the same person and the model returns someone who is not in any of them. Not badly rendered, not distorted — a plausible, well-made face that belongs to a person who does not exist. Everyone who has worked with multi-reference generation has seen this, and the usual response is to add more references, which makes it worse.

    This is a fusion error. The model is not failing to lock an identity; it is locking one, and the one it locked is the average of the accounts you gave it. Six photos of one person under six different lighting setups, at six crops, at six ages, are not six pieces of evidence for the same face. They are six descriptions of slightly different faces, and averaging is a completely reasonable thing to do with a set like that.

    Reference hygiene beats reference count

    On multi-reference work, hygiene matters more than count, and it is not close: uniform lighting, a consistent crop, consistent age and hair and makeup, adequate resolution. Three clean references beat ten mixed ones.

    The slot count is not your constraint anyway. In Versely's catalog, 38 active image models accept reference images, with caps running from a single reference up to sixteen. Nano Banana 2 takes fourteen at eight credits a call. Kling Image O1 takes ten at three. Qwen Image Edit 2511 takes six at three. Midjourney V7 takes five. At the other end, Flux Kontext works from a single reference and does it well.

    Almost nobody fills the high-capacity slots with genuinely distinct, clean information; the people who try usually fill them with near-duplicates, which adds noise and nothing else. How many slots a shot actually benefits from is its own question, worked through in how many reference slots your shot actually needs.

    This is a different failure from the one in garbage reference in. That piece covers a single weak reference producing a weak generation. This covers several individually-fine references producing a wrong one, because they disagree.

    What the model is doing with a mixed set

    Think of each reference as a claim about the subject. When the claims agree, the overlap sharpens: this jawline, this eye spacing, this hairline, confirmed six times. When they disagree, there is no mechanism that says "one of these is the real one." There is only the union, and the model resolves it by finding a face consistent with as much of the evidence as possible. That face is nobody.

    Lighting is the worst offender because it changes apparent geometry. A face lit flat and frontally has a wider-looking jaw and a flatter nose than the same face lit from forty-five degrees with the shadow falling across one cheek. Feed both and the model has two jaw widths to reconcile. It will reconcile them.

    Crop is the second offender, for a subtler reason: it changes what fraction of the reference is face. A tight head-and-shoulders crop and a full-body shot put very different amounts of identity signal into the same slot, so the tight crop effectively gets more weight. Mixing crops is mixing weights without meaning to.

    The third is styling drift over time. Photos from different years are photos of a face that has changed, plus a hairstyle that has changed, plus makeup that has changed — three more variables to average.

    The hard reference spec

    Apply this to every image in the set, without exception. An image that fails any row comes out of the set rather than getting waved through.

    Property Spec What breaks without it
    Lighting One setup across the whole set. Soft, broad, roughly frontal. No coloured gels, no hard side key, no strong backlight Apparent bone structure changes between refs; the model averages the difference into a new face
    Crop Torso-up on every image. Subject fills most of the frame Uneven identity weight per slot
    Angle Cover distinct angles deliberately: one front, one three-quarter, one profile. Not three fronts Duplicate angles waste slots; missing angles leave the model guessing at the sides of the head
    Styling One hairstyle, one makeup state, one apparent age, across all refs The averaged output splits the difference on all three
    Wardrobe Consistent if wardrobe is part of the identity; neutral and plain if it is not Clothing bleeds into the character definition
    Resolution At least 1024px on the short edge, and sharp when you view the crop at 100% Soft refs contribute mush, and mush averages toward generic
    Compression Original files. No screenshots, no re-saved-twice JPEGs, no social-platform downloads Compression artefacts get read as texture
    Background Plain and consistent is safest. Varied is tolerable if everything else holds Busy backgrounds compete for the model's attention with the subject
    Expression Neutral for the identity refs. Expression variants belong in a separate set Expression bleeds into the locked face

    The rows most often broken in practice are lighting and resolution, and they break for the same reason: the images were gathered rather than made. A folder assembled from whatever existed will fail the lighting row almost by definition.

    Building a set that holds

    The fix is to stop collecting references and start producing them.

    1. Generate or shoot one hero image. One frame you are genuinely happy with, lit the way the spec asks — soft, broad, roughly frontal, neutral expression, torso-up. Everything downstream derives from this. Spend real effort here; it is the cheapest place to spend it.
    2. Derive the angles from the hero. Three-quarter and profile generated from the hero rather than shot independently. Derived angles inherit the hero's lighting and styling automatically, which is the whole point.
    3. Check each derived frame against the hero before you keep it. Same hairline, same eye spacing, same jaw. Reject anything that drifted. This is the step people skip, and skipping it puts a subtly different face into the set that then contaminates everything.
    4. Upscale anything below the resolution floor. Upscaling an image's resolution before it enters the set is mechanical insurance, and cheap next to a generation you have to discard.
    5. Keep the set small and clean. Three to five images that all pass beats a dozen that mostly do.
    6. Register it once. Once a set holds, promote it out of your file system. Reusable characters and products registers the set under a key so every later prompt references the same images rather than whatever you happened to drag in that day. This is also what stops the set quietly drifting as people add "one more good photo" to it over months.

    When you cannot reshoot

    Sometimes the subject is a real person, the photos are what exist, and none of this is available. Three moves, in order of preference:

    Cut rather than average. Pick the single best-lit, best-resolution image and use it alone against a single-reference model rather than feeding five mixed ones to a multi-reference model. One clean claim beats five contradictory ones, and this is genuinely the most common right answer.

    Normalise before you submit. Crop everything to the same framing. Upscale everything below the floor. Drop anything with coloured or hard-directional light. You are not making the set perfect, you are removing the images that are actively pulling the average away.

    Normalise the light itself. If a reference is right in every way except that it is lit hard from one side, a relighting or edit pass to bring it toward soft frontal light turns a rejected image into a usable one. This is worth doing when the image contributes an angle you otherwise do not have.

    What none of these fix is a set where the subject genuinely looks different across the photos. If the hair is different and the age is different, no amount of normalisation makes them one person, and you have to choose which version of the subject you are locking. Reference sets are one of four distinct ways to hold a look across generations, and they fail differently from the others — four consistency mechanisms and when each fails covers when a reference set is the wrong tool entirely.

    FAQ

    Should the same reference set be used for images and video?

    The spec is the same, but video reference slots are scarcer and a bad generation costs more, so the discipline matters more. Build one clean set and use it for stills and for reference-to-video work alike. What you should not do is build a video set by pulling frames out of a previous video: generated frames carry accumulated drift, and feeding them back in compounds it.

    How do I tell a fusion error from ordinary identity drift?

    Fusion errors are stable and drift is not. Five images, five slightly different faces is drift, and the fix is a stronger anchor. Five images and the same wrong face every time is fusion — the model locked an identity successfully, it just locked the average of your set. Stable-but-wrong points at the reference set rather than the prompt.

    Does adding a text description of the face help?

    Marginally, and it can hurt. Text and references compete: if the prompt says "sharp cheekbones" and the references show softer ones, you have added a contradictory claim. Keep identity description out of the prompt when references are carrying identity, and spend the prompt on what references cannot express — action, camera, setting, light.

    What about product shots rather than faces?

    Same spec, one addition: material behaviour is part of a product's identity, and material reads differently under different light. A glossy bottle under a softbox and under a bare bulb looks like two finishes. Hold the lighting row even more strictly for products, and shoot the label straight on in at least one reference, since packaging text degrades first.