Guides

    Wan V2.6 Text to Video: write the sound into the prompt (5s, 10s, 15s, 1080p, 15cr)

    Wan V2.6 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    Versely Team4 min read

    Wan V2.6 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    Wan V2.6 Text to Video is text-to-video generation with WAN V2.6. Category is text-to-video only. audio is true. Durations are 5s, 10s, and 15s — three beats, not a free slider. max_output_resolution is 1080p. Fifteen credits is the catalog figure. requires_image is false: you describe the shot, including the sound.

    Three lengths. Write the room.

    5s is a hook. 10s is a beat. 15s is the long take this row will give you. There is no 6s, no 8s, no 12s. If your cut is 8 seconds, you are either trimming a 10s file or you are on a different model. Plan the prompt for one of the three enums.

    Because audio is native, the 5 / 10 / 15 seconds include a soundtrack. Dialogue, room tone, footsteps, a score the model invented because you did not forbid it. Leaving sound unspecified is how a skincare ad arrives with a saxophone. Write the line in quotes if someone speaks. Write "no music, room tone only" if they do not. Write "silence, mute plate" if you intend to score later.

    The AI video generator is the Versely door. Wan's roster and best Wan model are the family map. This page is only V2.6 text-to-video: native audio, 5s / 10s / 15s, 1080p, 15 credits.

    Cinematic, documentary, or abstract is a field

    Styles on this row are Cinematic, Documentary, and Abstract. That is an enum, not a vibe word in the prose. Pick one. Stacking "cinematic documentary abstract" in the prompt while leaving the field blank is how the style and the sentence fight.

    Aspect: 16:9, 4:3, 1:1, 3:4, 9:16. Quality settings: 720p and 1080p. max_output_resolution is 1080p. Vertical talking-adjacent shots still need the sound written; a 9:16 file with mystery music is not "fine for Reels."

    This is not image-to-video. There is no required still. If a specific pack or face must appear, lock it on a stills row first — Wan Preview 2.5 Text to Image is the video family's frame generator — then change rows. Do not describe the SKU in adjectives and hope V2.6 invents the label you would print.

    Fifteen credits is not a reason to skip the soundtrack sentence

    The catalog credits value is 15. Native audio is already in that call. You do not save credits by omitting sound; you spend them on a band you will duck. Native audio versus TTS is the rule for when a locked brand voice should replace this: not on the same mouth, and not as a silent prompt "so we can add VO later" on a talking shot.

    If the mouth is the shot and the voice must be the same person next week, this row is a one-off plausible speaker, not a series lock. TTS plus lipsync is the series tool. Wan V2.6 is the in-world tool: one pass, picture and sound agreeing by construction.

    Best model with native audio is the wider list. Best text-to-video is the category rank. Write the sound. Then generate.

    FAQ

    What happens if I do not mention audio in the prompt?

    You still get a soundtrack. audio is true. The model will choose one. That is the failure this page exists to prevent. Specify dialogue, ambience, music, or silence.

    Can I generate 8 seconds?

    Not on this slug. supports_durations is 5s, 10s, and 15s. Trim a 10s take or pick a model that lists 8s. Do not prompt "eight seconds" and expect the enum to move.

    Is this image-to-video?

    No. Category is text-to-video. requires_image is false. A locked still belongs on an image-to-video row, or on Wan's stills sibling, not as a hope inside this prompt.

    What do 15 credits buy?

    One Wan V2.6 Text to Video generate as listed: a 5s, 10s, or 15s clip, up to 1080p, native audio. It does not buy a mute-only mode you forgot to request. Prompt the sound you want.