Comparisons

    How many reference slots your shot actually needs

    Single-image, three-slot and nine-slot reference models compared: which shots improve past three references, which just get slower, and a shot-type table.

    Versely Team9 min read

    Nineteen video models in the catalog accept reference files, and their image-slot caps run from three to ten. Ten of those nineteen sit at exactly nine slots. Almost nobody fills nine, and the people who do are usually filling them with the same photo shot from slightly different distances, which is worse than leaving them empty.

    Slot count is a capacity spec. It tells you how many files the dispatch will accept. It says nothing about how many the model will meaningfully weight, and nothing about whether your particular shot has that many distinct things worth locking. Here is how the classes actually differ and where the returns stop.

    Three classes, not a spectrum

    The single-image class. Sixty-eight of the 146 video models require an image input, and only two of the 52 image-to-video entries accept more than one reference file. So the default in this catalog is one frame: you supply the opening image and the model animates from it. That is not a reference relationship at all, which is the distinction most people miss. The frame is the literal first frame, so identity is guaranteed at t=0 and decays from there. Image-to-video and reference-to-video are different operations, and the failure modes are opposite: image-to-video starts perfect and drifts, reference-to-video starts approximate and holds.

    The three-to-five class. VEO 3.1 reference-to-video accepts three images and generates 8-second clips. Kling O3 accepts four across its standard and pro tiers, at up to 15 seconds. Wan 2.7 reference-to-video takes five images and, separately, up to five video clips. This band is built for one subject plus context: a face, a garment, a location. The constraint is honest and it forces a useful discipline, because with three slots you have to decide what actually matters.

    The nine-plus class. Happy Horse reference-to-video takes nine images across both its 1.0 and 1.1 entries. Seedance 2.0 takes nine images plus three video references and three audio references. MiniMax H3's reference mode carries the same nine-plus-three-plus-three shape, and its catalog description spells out the mechanism: up to nine images for subject and style, three video clips for motion, three audio clips, each cited in the prompt by order. Kling O1's reference mode and Gemini Omni Flash's reference mode both go to ten images.

    The interesting thing about the top class is not the image count. It is that the nine-slot models are the ones that also accept non-image references. Nine images is a lot of the same kind of information. Nine images plus a motion clip plus a voice sample is three kinds.

    What a slot is actually worth

    A reference slot is worth something in proportion to the information it adds that no other slot carries. That sounds obvious and it is routinely violated, because the natural instinct when a model offers nine slots is to upload nine frames from the same photoshoot.

    Ranked by how much a marginal slot typically buys:

    1. A new angle on the same subject. Front, three-quarter, profile. This is the highest-value second and third slot for any character or product work, because it is the only way the model learns what the back of the head or the side of the bottle looks like.
    2. A new lighting condition. A subject shot in daylight and again under warm interior light tells the model what is identity and what is illumination. Without it, the model can bake your key light into the character.
    3. A different kind of thing entirely. The set, the product, the garment. A slot that holds the location is doing work no portrait slot can do.
    4. A motion reference. On the models that accept video, one clip that carries the movement you want is worth more than four extra stills, because motion is the thing stills cannot express.
    5. Another near-identical frame. Close to zero, and occasionally negative. Redundant angles concentrate whatever is idiosyncratic about that one pose, which is how you get a character who can only stand one way.

    The practical rule that follows: fill slots with categories, not copies. If you cannot say in one word what a slot is for that the other slots are not, do not fill it. Reference quality matters more than reference count, and a bad reference actively degrades the output rather than being ignored, which is the input rule most people learn the expensive way.

    The addressing problem

    Capacity without addressing is a trap. A model that accepts nine images and gives you no way to say which one is the character and which one is the set will blend them, and blending is exactly what you were trying to prevent.

    The models worth using at high slot counts are the ones with a citation convention. MiniMax H3's reference mode expects each reference to be cited in the prompt by its order, so the prompt does the assignment rather than leaving it to inference. Happy Horse's reference workflow uses in-prompt character tags, so a two-character scene can be written with each character bound to a specific upload.

    The test for whether a model actually supports high slot counts is not the number on the spec sheet. It is whether the prompt syntax gives you a handle on each slot. If it does not, treat the model as a three-slot model regardless of what it accepts, because past three undifferentiated inputs you have lost control of what is being locked.

    Shot type to reference count

    Shot Slots that earn their place What goes in them Past this point
    Talking head, one person, one set 1–2 Face front, face three-quarter Extra portraits add nothing
    Product hero, exact SKU 2–3 Front, angled, label detail More angles help only if the shot rotates
    Person holding product 3 Face, product front, product in hand Fourth slot usually redundant
    Character across a campaign 3–4 Two angles, one lighting variant, one full-body Returns flatten hard at five
    Character in a branded environment 5–6 Character set, plus set, plus product Only worth it if the model addresses slots
    Motion matched to existing footage 2–3 images + 1 video Subject stills, one clip carrying the move Second motion clip rarely helps
    Voice or timbre matched Images + 1 audio Subject stills, one clean audio sample Longer sample beats more samples

    The pattern across that table: three is where most shots stop improving, and the exceptions are shots with more than one kind of thing to lock. A character in a branded set with a specific product is genuinely three subjects and reasonably needs six slots. A person talking to camera is one subject and needs two.

    Finding your own diminishing point

    The comparison takes one session and it is the only way to know where your shot type lands.

    1. Fix everything except reference count. Same prompt, same seed where the model exposes one, same duration, same aspect ratio.
    2. Run three versions: one reference, three references, and the model's maximum.
    3. Judge on identity only. Not which clip you like. Whether the subject in frame is the subject in your references, at the start, middle and end of the clip.
    4. Look for the flip. Frequently the maximum-reference version is worse, with a subject who looks like an average of the uploads rather than any one of them. That flip is the answer.
    5. Write the number down next to the model and the shot type. It will be stable for that combination until the model version changes.

    If one and three look identical, you are on a shot that does not need references at all and an image-to-video model with a single strong opening frame will be faster and cheaper. If three beats one clearly and the maximum beats three, you have a shot that justifies the high-slot class, and the reference-to-video comparison is where to pick between the models in it.

    FAQ

    Should I always fill every available slot?

    No. Fill the slots you can justify by category, then stop. Empty slots cost nothing; redundant slots cost upload time, and on a model that averages its inputs they cost identity. The clearest sign you have overfilled is a subject who looks plausible but not like anyone in your reference set.

    Do video references count against the image slot count?

    On the models that accept both, they are separate pools. Seedance 2.0 and MiniMax H3's reference mode each expose nine image slots and three video slots and three audio slots as distinct budgets, so a motion clip does not cost you an image. Wan 2.7 splits the same way, with five images and five video clips. Check the model entry rather than assuming, because a single combined pool would change how you allocate.

    Is a three-slot model worse than a nine-slot model?

    Not for most shots. VEO 3.1's reference mode caps at three images and eight seconds, and for a single-subject shot that is not a limitation you will feel. The nine-slot models earn their extra capacity on multi-subject scenes and on jobs where motion or voice references matter. Picking the highest slot count available for a talking-head shot is paying for headroom you will not use.

    What about consistency across many separate clips?

    Reference count helps within a clip. Across clips you also need the references themselves to be stable, meaning the same files in the same order every time, and ideally the same seed family. Changing one reference between shots is enough to shift a face subtly, and subtle is worse than obvious because it survives review and shows up in the cut. Reference stacking for consistency covers the discipline for campaign-length work.

    Start from the reference-to-video models in the catalog, run the one-versus-three-versus-max test on your actual subject, and let the result pick the class rather than the spec sheet.