LTX 2 Audio to Video: the mouth pass after the picture exists (20cr)
LTX 2 Audio to Video wants a plate and a track. It is not how you invent the scene.
LTX 2 Audio to Video wants a plate and a track. It is not how you invent the scene. Content type is lipsync. Category is audio-to-lipsync. audio is true. Credits 20, flat. Durations empty. Max output empty. requires_image is false. The description: generate videos from audio input with synchronized visuals and motion. Features: audio_to_video, audio_visualization, music_video. The track is the brief. This row does not design the set, the lighting, or the SKU.
The lipsync tool is the launcher. The LTX roster is the family. Best LTX model and best lipsync model are the rankings. This page is only the audio-to-video slug at 20 credits.
The track is the brief
Audio-to-lipsync means the sound file is the input that matters. Length of the result follows the track, which is why supports_durations is empty — you do not pick 5s or 8s on this row. You bring a VO, a song, a line. The model syncs visuals and motion to that.
If you do not have a track yet, you are early. Write and record (or clone) the audio first. Then come here. Inventing a scene in a text-to-video prompt and hoping the mouth later matches a VO you have not made is the expensive order.
requires_image is false on the snapshot, so a plate is not a hard catalog gate. Bring one anyway when identity matters. A selfie or a locked still is how the mouth belongs to someone you already approved. Without a plate, you are asking the row to invent a face and sync it. That is two jobs. This slug is the second one.
Twenty credits is a mouth pass, not a world
20 credits, billing type flat, min and max both 20. One call. It is not a 4K cinematic generate and not a 15-second text-to-video hold. It is a lipsync job at 20 credits.
Features include audio visualization and music video. Those are still driven by the track. A waveform look or a music-video motion pass is not a reason to skip locking the performer. Lock the still on a 1–3 credit image row. Then spend 20.
Do not use this as a substitute for text-to-video. LTX has video slugs for that. This slug's content type is lipsync.
Do not invent the set here
World, wardrobe, product, camera — those are stills and video rows. This row takes a finished picture problem (who is on camera, or a visualization of a mix) and a finished audio problem (what they say, or what plays) and returns synchronized motion.
If the still is wrong, 20 credits will animate the wrong still. If the track is a scratch VO with the wrong line, 20 credits will lip-sync the wrong line. Fix inputs. Then generate.
HeyGen's photo-avatar row is a different contract (selfie + voice, 50 credits). Sync 2.0 is video-to-lipsync with named sync modes. LTX 2 Audio to Video is the 20-credit audio-to-lipsync cell with visualization and music-video tags. Pick the row that matches the inputs you actually have.
FAQ
Is this text-to-video?
No. Content type is lipsync. Category is audio-to-lipsync. The input that defines the job is audio.
Do I have to upload a still?
The catalog does not require an image. Bring a plate when you need a specific face. Without one, the row still wants a track.
Why are there no 5s / 8s options?
supports_durations is empty. Length follows the audio you supply, not a clip-length dropdown.
What does 20 credits cover?
One flat generate on this slug as listed: audio in, synchronized visuals and motion out. It does not buy a scene design pass. Design the scene on a stills or video row first.