LTX 2 Text to Video Pro: write the sound into the prompt (6s, 8s, 10s, 4K, 6cr)
LTX 2 Text to Video Pro has native audio. A silent prompt still gets a soundtrack you did not choose.
LTX 2 Text to Video Pro has native audio. A silent prompt still gets a soundtrack you did not choose.
LTX 2 Text to Video Pro is professional text-to-video with LTXV2. Content type is video. Category is text-to-video. requires_image is false. audio is true. Durations: 6s, 8s, 10s. Max output: 4K. Quality chips: 1080p, 1440p, 2160p. Frame rates: 24fps, 30fps. Motion levels: Medium, High. Display price: 6 credits. Styles: Cinematic, Documentary, Abstract. The mix is on. The still is not required. The length list is shorter than rows that go to 15s or 30s — write the cue to 6, 8, or 10.
Native audio at 4K is the row
Native audio means the 4K take includes a soundtrack the model wrote. A prompt that only describes picture is still a mix brief. You will get room, bed, Foley, or a line you did not spec. Write who speaks, what they say, whether the bed is quiet, whether the room is dead. If you needed silence, say silence. The 6-credit generate produces audio either way.
The door is the AI video generator. Ranked: best model with audio, best 4K AI video model, best LTX model. Family: the LTX roster. Professional text-to-video is the description. Professional includes the mix.
Six, eight, or ten seconds — not fifteen
Duration chips are 6s, 8s, 10s. There is no 15s, no 20s, no 30s on this slug. A silent-looking 10s prompt is ten seconds of leftover score at up to 4K. Write the cue to the length you picked. If you need a longer hold, that is another row or an extend pass, not a prompt trick here.
4K is the ceiling (2160p on the quality list, with 1080p and 1440p below). 6 credits is the display price. That combination — native audio, 4K, 6/8/10s, 6 credits — is why people treat it as a disposable preview. The preview has a mix. The mix may already contain speech that fights the VO you planned to add.
24fps and 30fps are real chips. Pick one on purpose. Medium and High motion are motion levels on a text-to-video take, not a reason to skip writing the sound.
Silence in the prompt is still a mix
Documentary style still gets audio. Abstract still gets audio. Cinematic still gets audio. Style is a picture register. The audio flag does not follow it.
No image required. If you have a locked plate, confirm you wanted text-to-video. This Pro slug will invent a 4K scene from text, with a soundtrack, in 6–10s. It will not honor a product still you forgot to attach because the row does not require one.
Unmute the 4K take before you caption it
Play the take muted, then unmuted. If the unmuted version contains a riser, a line, or a room you cannot ship, the prompt was incomplete. Put the sound in the next 6-credit run. Do not "fix it in captions." Captions will transcribe the line you did not write.
If the job is a silent product turn off a locked still, this is the wrong family. If the job is a 6s, 8s, or 10s 4K scene with diegetic sound for 6 credits, this is the row — and the prompt includes the soundtrack.
FAQ
Does LTX 2 Text to Video Pro always return sound?
audio is true. Treat native audio as on. A picture-only prompt still gets a soundtrack you did not choose. Write the cue, or write silence.
What lengths and resolution?
6s, 8s, and 10s. Max output 4K (1080p / 1440p / 2160p chips). 6 credits display. No image required. Text-to-video.
Can I use this as a silent 4K preview?
You can mute in the editor. The file you paid for still has a mix. Prompt the sound you want, or pick a silent row.
Is this the same as a 15s or 30s audio-on model?
No. This slug stops at 10s. Longer native-audio rows are other slugs. Do not prompt LTXV2 Pro to "hold for 30."