Guides

    Seedance 2.0: the stack is the brief (4s, 5s, 6s, 1080p, 69cr)

    Seedance 2.0 is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.

    Versely Team4 min read

    Seedance 2.0 is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.

    Seedance 2.0 is ByteDance's cinematic video row: Hollywood-grade footage from a prompt, native audio-visual synchronization, director-level camera and lighting control, exceptional motion stability. Categories are text-to-video and reference-to-video. requires_image is false, which is how teams skip the stack and then spend 69 credits discovering the face is a cousin. The claim of this page is the second category. Upload the stills that must appear.

    4s is a test. 15s is a bill.

    supports_durations is 4s through 15s, every second. The title on this page names 4s, 5s, and 6s because those are the honest first calls. The slider going to 15s is not a reason to start there. A 15-second identity miss costs the same kind of credit as a 4-second one, only more of them.

    max_output_resolution is 1080p. Quality settings on the row include 480p, 720p, 1080p, and 4k. Audio is on. Native audio-visual sync is in the catalog description, not an add-on. A silent prompt still gets a soundtrack you did not choose. Write the room: dialogue or none, footsteps or none, score or none.

    Aspect includes 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive and auto. Pick the cut. Do not leave auto on a 9:16 brief and crop the extra later; you paid 69 credits for pixels you will throw away.

    The stack is the brief. The prompt is blocking.

    This is not image-to-video. Frame one is not automatically your upload. Reference-to-video recasts the people and products you showed it into a new shot. If you needed the kitchen you photographed to stay that kitchen, you are on the wrong category. The split is image-to-video vs reference.

    Seedance 2.0 will run from text alone. That is the text-to-video category, and it is how a 69-credit cinematic clip becomes a casting call you did not hold. For identity, upload the stills. Do not describe the face in adjectives and hope 69 credits remember it. How Seedance family prompts point at files is a separate routing note — @-references versus dedicated R2V — and it is not a reason to skip the upload.

    Register the set once with reusable characters and products. Clip nine should pull the same keys as clip one. Reference hygiene still applies: same crop family, same light, no mixing a studio packshot with a phone grab.

    69 credits is one generate, not a dare

    The catalog credits figure is 69. That is a cinematic 1080p call with audio on, not a coupon to max duration "because we already opened the row." Lock stills on a text to image pass until the face and the label would print. Then spend Seedance once, at the shortest duration that holds the beat.

    Camera and lighting belong in the prompt because the description names director-level control. Name the move. Name the key. Do not also re-describe hair colour the stills already answered. The stack and the sentence should not fight.

    ByteDance's other rows are on the ByteDance provider page. The category ranking is best reference-to-video. This page is only Seedance 2.0: 4s–15s, 1080p, native audio, 69 credits, stills first, slider last.

    If the job is "this person, this bottle, new scene," you are in the right place. If the job is "animate this exact frame," leave. Image-to-video is the other door.

    FAQ

    Does Seedance 2.0 require an image?

    No. requires_image is false. It will generate from text. If a specific face or SKU must appear, treat that as a reason to upload anyway. Reference-to-video is the job this page is arguing for. Text-only is how the 69-credit take becomes a recast.

    Why not always run 15 seconds?

    Because the catalog lets you. 4s, 5s, and 6s exist so a beat can be a beat. A 15-second hold of a wrong identity is not "more value." It is a longer wrong file with native audio you now have to throw away.

    Is the output silent if I do not mention sound?

    No. audio is true. Native audio-visual synchronization is in the description. A mute prompt is not a mute clip. Write the soundtrack you want, including silence if you need a plate to score later.

    Is this the same as Veo reference-to-video?

    Same job family, different row. Seedance 2.0 is ByteDance, 4s–15s, 1080p, 69 credits, and it also lists text-to-video. Veo's reference row is a different slug with its own duration cap. Do not copy a Veo stack budget onto this slider.