How to spend H3's nine reference slots
MiniMax H3 takes nine reference images, three videos and three audio clips. A working slot budget for a character-plus-product series, and what to drop first.
Nine reference images sounds like abundance until you sit down with a real brief. One presenter, one product, a room they're standing in, and a look you need to hold across six shots. Suddenly nine is tight, and the way most people fill it — nine slightly different photos of the same face — wastes seven of them.
MiniMax H3's reference-to-video mode takes up to nine images for subject and style, three video clips for motion, and three audio clips, with each reference cited in the prompt by its order in the set. That last detail is the one that changes how you should load them, and it's the one people discover on shot four of a series when everything stops matching.
Here's a slot budget that survives a real character-plus-product campaign.
The rule that decides every slot
A reference slot is expensive. A sentence is free. So:
Never spend a slot on something a sentence can specify.
"Warm tungsten key light with soft falloff" does not need a reference image. "Shot on a long lens, background compressed" does not need one. "She smiles" does not need one. Those are all promptable in words, reliably, at zero slot cost.
What words cannot specify is identity. A particular face. A particular label with particular type on it. A particular silhouette in a particular jacket. Those are what slots are for, and if you apply this one rule you will find you have more slots than you thought, not fewer.
A slot budget for a character-plus-product series
Nine slots, allocated in three blocks. This is the layout I'd load for a six-shot series with one presenter and one product.
| Slot | Reference | What it locks |
|---|---|---|
| 1 | Character, front on, neutral expression, even light | Face structure. The anchor frame. |
| 2 | Character, three-quarter, same lighting | Depth and volume of the face |
| 3 | Character, full body, in the wardrobe for the series | Build, silhouette, clothing |
| 4 | Character, mid-action or mid-speech | How the face behaves when it isn't posed |
| 5 | Product, hero angle, neutral ground | Shape and proportion |
| 6 | Product, tight crop on label or key detail | Typography, finish, materials |
| 7 | Product held in a hand | True scale, which is the thing models get wrong |
| 8 | Environment plate, empty of people | Set, architecture, palette |
| 9 | Style plate: a frame with the exact grade you want | Look, when words aren't landing it |
Four for the character, three for the product, two for the world. That split reflects where generative failure actually happens: faces drift most, product proportions second, environments least, because a room is more forgiving than a jawline.
Three notes that matter more than the layout itself:
Slot 7 is underrated. A model that has never seen your bottle next to a hand will invent a size, and the invented size is usually wrong in a way that reads as fake before the viewer can say why. One in-hand frame fixes it.
Slots 1 and 2 must share lighting. If your front-on frame is studio-lit and your three-quarter is a phone snap in a kitchen, you have given the model two different people and asked it to average them. Consistent lighting across the identity frames is worth more than a third angle.
Fix the order and never break it. Because references are cited by their position in the set, the order is part of your prompt. Load the same nine in the same order for every shot. Reordering them between shots silently rewires every reference in your prompt text, and the symptom — shot four suddenly holds the wrong thing — looks like model instability rather than what it is. Write the order down before the first generation and treat it as frozen.
A prompt that uses the set looks like this:
The woman in images one and two walks into the room from image eight, holding the bottle from images five and six. She turns the label toward camera around the six-second mark, then sets it down. Match the grade of image nine: warm tungsten key, soft shadow falloff, shallow depth of field. Handheld, slight drift, no cuts.
Note what is doing the work. The references handle identity. The sentence handles behaviour, timing and camera. That division is the whole technique.
The three video slots and the three audio slots
The video slots take motion, not content. One well-chosen clip of a camera move is worth more than three of vaguely similar footage. My default allocation:
- Camera behaviour. The push-in, the orbit, the handheld drift.
- Subject action. How the person moves, if describing it in words fails.
- Leave empty on the first pass. Add only if shot one comes back with a motion problem you can point at.
The audio slots need one thing said plainly first: H3 returns picture only. Nothing you load into those three slots comes back in the render. They are performance inputs — audio the model uses to drive what the picture does, not a mix it hands back. Score and mix H3 footage in post the way you would camera-original material; the H3 review covers that finishing pass.
So allocate them for performance, not sound design:
- Speech you want the mouth to follow, if the shot has someone talking.
- A rhythm or pace reference, if the action needs to move to a tempo.
- Leave empty. Music belongs in the edit, timed against the cut.
Under-filling is a legitimate strategy. Every reference is a constraint, and a heavily constrained generation has less room to produce the good accident you were hoping for.
Drop order when the set is full
You will hit nine. A second character joins the series, or a second SKU. Here is the order to cut, most droppable first, with the reason.
- Slot 9, the style plate. Grade is the most promptable thing on the list. "Warm tungsten, soft falloff, low contrast, shallow depth" gets you most of the way, and the character frames carry lighting information anyway.
- Slot 8, the environment plate. Drop it if the setting is generic — an office, a kitchen, a street. Keep it if the location is specific and recognisable, because a real place is not describable.
- Slot 4, the character in action. The identity frames already carry the face. This slot buys expression range, which is a nice-to-have on a product beat and a must-have only on dialogue-heavy work.
- Slot 7, the in-hand scale shot. Only drop this if the product has an obvious, unambiguous real-world size. A phone, a mug, a book will survive. A cosmetic bottle, a supplement tub or anything with unusual proportions will not.
Never drop slots 1, 2 and 5. Front-on character, three-quarter character, product hero. That trio is the identity floor, and everything below it produces a series that doesn't feel like a series. If you need a second character and you're already at nine, cut slots 9, 8 and 4 to free three, and give the new character front-on, three-quarter and full-body.
Tuning the set, and where the budget transfers
H3 bills per second of output, so duration is the lever, not resolution. Tune your reference set on 5-second generations, which is the floor, and only extend to the full 15 once the set holds identity across two or three consecutive shots. Testing a nine-slot arrangement at maximum duration is spending on footage you already know you're going to throw away.
Run the whole thing through the agent if you're iterating: one instruction can fan the same prompt and reference set across more than one model in a single request, which is the fastest way to find out whether your set is weak or the model is.
The 9 / 3 / 3 shape is not unique to H3. Seedance 2.0 and its fast reference tier take the same nine images, three videos and three audio clips, so a set built to this budget moves between them without rework. One difference worth planning around: Seedance generates synchronised audio with the picture, so on that family the audio slots steer something you actually get back. Same slot budget, different finishing pass. That portability is the real argument for building the set properly once rather than improvising per model.
The broader technique is just character consistency discipline applied to a hard slot ceiling: decide what only an image can say, spend slots on exactly that, and say everything else in words.
FAQ
How many references does MiniMax H3 accept?
Up to nine images, three video clips and three audio clips in a single generation, each cited in the prompt by its position in the set. Reference-to-video is one of three H3 modes in the catalog, alongside the text-to-video and image-to-video variants, all of which share the same durations and resolutions.
Should I always fill all nine image slots?
No. Every reference is a constraint, and over-constraining produces stiff, literal output. Fill the identity slots that words cannot cover — face, product, specific location — and leave the rest empty. Under-filled sets are frequently better, and they are certainly easier to keep consistent across a series.
Does the order of my reference images matter?
Yes, and this is the failure most people hit late. References are cited by position, so changing the order between shots rewires every reference in your prompt text. Fix the order before the first generation and keep it frozen for the whole series.
What should I cut first when I need a slot for something new?
The style plate, then the environment plate, then the character action frame, then the in-hand scale shot. Keep the front-on character frame, the three-quarter character frame and the product hero under all circumstances — those three are what make a set of clips read as one campaign.
Load a set and run it in the AI product video generator if the SKU is the hero and the presenter is the support act, or in the general video generator if the balance runs the other way.