Seedance 1.5 Pro Text to Video: write the sound into the prompt (4s, 5s, 6s, 1080p, 6cr)
Seedance 1.5 Pro Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.
Seedance 1.5 Pro Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.
ByteDance's professional text-to-video row lists 6 credits, 1080p, and durations from 4s through 12s in one-second steps. Aspects run 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and auto. Qualities are 480p, 720p, and 1080p. Catalog audio is true. The generate-audio control on this family defaults on. A prompt that only names choreography still comes back with a score, a whoosh, or a line.
Write the sound as part of the motion, because motion is what this model is for.
Choreography first, then the room
Seedance's family tip is movement patterns, choreography, and dynamic visual elements — not named camera hardware. Lead with what the body does for 4s, 5s, or 6s (or the longer 7–12s hold if the move actually lasts). "She steps left, the skirt follows, the second dancer enters from frame right." Then the sound that belongs to that move: floor creak, breath, no lyric, or the one spoken count-in.
A silent prompt on a dance-shaped model is how you get a music bed in a key you will hate at cut. Name it:
- Count, clap, or bare room.
- One instrument if the move is musical.
- Speech only when a mouth is in frame, and then the exact words.
The native audio versus TTS rule still holds. Seedance will invent a singer. It will not invent your singer. If the series voice is a clone, do not put a singing face in this generate.
1080p at 6 credits is the listed contract
Seedance 1.5 Pro Text to Video is the 1080p professional text-to-video card, not a preview of Seedance 2.0. Max output resolution is 1080p. The 6-credit figure is the catalog rate for this row. Pick 480p or 720p when the cut is a draft; pick 1080p when the plate is the plate.
Duration is a hard enum, 4 through 12. 4s, 5s, and 6s are the social-length windows most briefs actually need. Stretching a two-second gesture across 12 seconds is how the model pads. Compressing a 12-second phrase into 4s is how it smears. Size the choreography to the second you bought.
Aspect is a dropdown, including 21:9 and auto. Set it. "Cinematic widescreen" in prose does nothing if the control is still 9:16.
Do not treat silence as a savings
You are not saving 6 credits by omitting the soundtrack from the prompt. You are paying 6 credits for a track you did not art-direct. The AI video generator will still run the row. ByteDance's roster also holds later Seedance generations; those are different cards, different rates, different duration ceilings. This page is 4s–12s, 1080p, native audio, 6 credits.
If the shot has no mouth and you truly want mute, say "no speech, no music, only the sound of the move." Then listen. Native audio rows fill holes.
The wider roster sits on best text-to-video. Seedance 1.5 Pro is the motion-first, audio-on, 1080p member of that list — not a silent sketch model.
FAQ
Does Seedance 1.5 Pro generate audio even if I never mention sound?
Yes. Catalog audio is true, and this family's generate-audio default is on. A silent prompt still gets a soundtrack you did not choose. Write the room, or write "quiet, no speech."
Why start with 4s, 5s, and 6s instead of always running 12s?
Because the move has a real length. 12s is in the enum; it is for a phrase that lasts 12 seconds. Padding a short beat wastes the back half and still bills the longer clip.
Is 1080p the same as a 4K Seedance row?
No. This card's max output resolution is 1080p, with 480p and 720p as quality steps. Do not promise a 4K master from 1.5 Pro Text to Video.
Can I add a cloned founder voice on top of a Seedance take?
Not on a mouth that already spoke. Native audio already chose a voice. If you need the clone, keep faces from speaking in the generate, or pick TTS and a lipsync row instead of fighting this soundtrack.