Guides

    Flux 3 Text to Video: write the sound into the prompt (5s, 6s, 7s, 9cr)

    Flux 3 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    Versely Team4 min read

    Flux 3 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    Flux 3 Text to Video is Black Forest Labs' frontier video model — video with native audio directly from a text prompt, at up to 1080p, 5–20 seconds, across eight aspect ratios. Category is text-to-video. audio is true. Nine credits is the catalog figure, billed per second. requires_image is false. Features: native_audio, frontier, long_duration, 1080p.

    5s, 6s, 7s — then every second to 20

    supports_durations is 5s through 20s, every second. Wan V2.6 only offers 5 / 10 / 15. Kling Turbo starts at 3s and is silent. Flux 3 starts at 5s, can land on 6s or 7s, and will go to 20s. That is the long-duration feature. It is not a reason to start at 20s.

    A 20-second native-audio miss is a long miss with a band you did not pick. Write a 5s, 6s, or 7s beat until the sound and the picture agree. Then lengthen.

    Quality settings: 720p and 1080p. The description says up to 1080p. max_output_resolution is not a separate 4K field on this row. Do not prompt 4K and expect a 4K enum that is not here.

    The AI video generator is the door. Black Forest Labs is the video-side brand hub; Flux image models are the stills line. Best Flux model ranks the family. This page is only Flux 3 text-to-video: native audio, 5s–20s, 9 credits a second.

    Eight ratios, including 21:9

    Aspect: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16, and auto. Eight. Widescreen is a first-class option, not a crop of 16:9. Pick it. Leaving auto on a 9:16 brief is how you pay 9 credits a second for extra width you will throw away.

    Because audio is native, the prompt is a sound brief as well as a picture brief. Write the line in quotes if someone speaks. Write the room if they do not. Write "no music" if a saxophone would kill the ad. A silent prompt on this row is how a frontier model scores your clip for you.

    Native audio versus TTS is the rule when a locked brand voice should replace this. Not on the same mouth. Not as a mute prompt "so we can add VO later" on a talking shot. Flux 3 will still emit a soundtrack; you will duck it and keep the original lips.

    Nine credits a second is already the soundtrack

    The catalog credits value is 9, per second, matrix min 43 / max 170. You do not save 9 credits by omitting sound. You spend them on a bed you will replace. Specify silence only if you need a mute plate, and say so in the prompt rather than hoping frontier means quiet.

    Lock stills first if a SKU or face must match — Flux 2 Flash is the 1-credit 4K iterate on the image side — then change rows. Flux 3 text-to-video will invent a bottle from adjectives. That is allowed. It is also how identity retries land on the 9-credit meter.

    Best model with native audio is the wider list. Best text-to-video is the category. Write the sound. Then generate. 5s, 6s, or 7s until you mean 20.

    FAQ

    What happens if I do not mention sound?

    You still get a soundtrack. audio is true. Native audio from a text prompt is the catalog description. Specify dialogue, ambience, music, or silence.

    Can I generate 4 seconds?

    No. The enum starts at 5s and runs to 20s. 5s, 6s, and 7s are the short calls on this slug. Kling Turbo is the row that lists 3s and 4s, and it is silent.

    Is this the same as Flux 2 Flash with a duration?

    No. Flash is text-to-image, 1 credit, 4K, no audio. Flux 3 Text to Video is a Black Forest Labs video model, 9 credits a second, 5s–20s, native audio, up to 1080p.

    What do 9 credits buy?

    The catalog figure is 9, billed per second. You are buying one text-to-video generate with native audio, 5s–20s, 720p or 1080p, one of eight aspect ratios. Prompt the sound you want. Do not pay 9 credits a second for a surprise score.