Guides

    Seedance 2.5's reference budget: when fifty slots matter

    2.5 accepts far more multimodal references than 2.0. You only need that when identity and product must lock across a sequence.

    Versely Team6 min read

    Fifty slots is not a default. Seedance 2.5's reference-to-video endpoint will take up to 30 images, 10 videos and 10 audio files — fifty multimodal inputs, each addressable from the prompt. Seedance 2.0 topped out at 9 / 3 / 3. That jump is real. It is also a trap if you fill it because it exists.

    A single product still plus a character still is the usual job. Use the fat budget when identity and product have to lock across a sequence. Otherwise you are paying attention, and sometimes credits, to condition the model on noise.

    What the catalog actually publishes

    Seedance 2.5 is the text-to-video sibling: one prompt, one take up to 30 seconds. Seedance 2.5 reference-to-video is the endpoint with the budget. The description is the contract: generate from up to 50 multimodal references — images, videos and audio — and address them as @Image1, @Video1, @Audio1.

    That is a different API from image-to-video. I2V animates a frame you already like. R2V casts a subject into a scene you describe. If those two sentences still blur, image-to-video vs reference explained is the split; this page is only the budget.

    Seedance 2.0 R2V Seedance 2.5 R2V
    Image refs 9 30
    Video refs 3 10
    Audio refs 3 10
    Single-pass duration 4–15s 4–30s

    2.5 is "far more" in the literal sense. It is not "always more." Nine images was already a cast, a pack, and a room. Thirty is a shoot binder.

    The usual job is two stills

    Most commercial R2V work is one product and one person, sometimes a room.

    • Product still. Clean, labelled, the angle you would put on a PDP. This is identity, not lifestyle.
    • Character still. Face, wardrobe, palette. One frame you would actually recast from.
    • Optional third. The set, if the room is a brand asset rather than a prompt.

    That stack would have fit in 2.0 with room to spare. It still fits. The reason to be on 2.5 for this job is the 30-second pass and the tighter lock, not the extra 41 empty slots.

    Filling those slots with "a few more angles, a moodboard, last month's ad, a competitor's still, a voice memo" is how identity gets averaged. Extra references are extra votes. Weak votes do not help a pack shot.

    Default prompt discipline: name the two or three files you mean (@Image1 is the bottle, @Image2 is the presenter), describe the new action and the camera, and stop. Do not re-describe the label the still already shows.

    When fifty slots are the product

    The fat budget earns its keep on a sequence, not a single hero.

    A cast. Several characters who have to stay themselves across cuts. Each face is a still. Each wardrobe is a still. You are not making a music-video extra; you are preventing the brunette from becoming a blonde at second 18.

    A product system. Bottle, box, bag, colourways, a logo lockup that marketing will reject if it warps. One still per SKU that has to appear, not a contact sheet of every SKU you sell.

    A set bible. Hero location plus the two connecting spaces the 30-second take will walk through. Video references help here for motion and blocking, not as a way to smuggle a whole edit into the generate.

    A voice or bed you must match. Audio refs exist so the take can lock to a line-read or a bed you already cleared. They are not a substitute for "make it sound expensive."

    If you cannot point at the sequence and say which identity would drift without the extra files, you do not need the extra files. Generate on the two-still stack. Save the binder for the episode, the launch set, or the three-shot ad where the same bottle has to survive a kitchen, a bag, and a close-up.

    Two ways to waste the budget

    Treating fifty as a quality slider. More references will not make a bad still good. A cluttered lifestyle photo of the product in a messy kitchen is a worse conditioner than one pack shot. Clean the still first.

    Mixing I2V and R2V on the same shot. I2V already pinned a frame. Stacking a reference pile on top of that contract is how you fight yourself. Pick the endpoint. The 2.5 family exposes both; the shot should use one.

    A third, quieter waste: uploading video refs as if they were a motion-control clip. A driving performance is a different job with a driving-clip input. Seedance's video slots are identity and motion hints inside one generate, not "dance like this" control.

    FAQ

    Do I need all fifty references to use Seedance 2.5?

    No. The ceiling is a ceiling. The usual commercial job is a product still and a character still. Use more only when a sequence would otherwise lose a specific identity.

    Why not always max out the slots "just in case"?

    Because unused slots are harmless and filled weak slots are not. Extra images are extra votes. A moodboard, a competitor frame, and three redundant angles will pull the take toward an average of all of them.

    Is this the same as attaching files with @ in a text-to-video prompt?

    Related family, different endpoint. Seedance 2.x @-files in one generate and the dedicated R2V catalog endpoint are not a stack to combine on one shot. Pick the endpoint the catalog exposes for the model. This post is the budget rule for the R2V row.

    When should I stay on Seedance 2.0?

    When the take is 15 seconds or under and nine images already cover the cast. 2.5 is the 30-second pass and the larger binder. It is not a mandatory upgrade for a two-still product clip.


    The new number is 50. The operating number is 2, sometimes 6, occasionally a binder. Fill the binder when the sequence has names. Leave it empty when the brief is a bottle and a person.