LTX 2.3 lipsync is a 20-credit mouth
LTX 2.3 lipsync is 20 credits at 1080p. Use it when the LTX picture already exists and you only need the mouth to follow a track.
Every guide, comparison and workflow we’ve published on AI Lipsync.
36 articles — page 1 of 2
LTX 2.3 lipsync is 20 credits at 1080p. Use it when the LTX picture already exists and you only need the mouth to follow a track.
Fabric is 40 credits at 720p. Avatar Pro is 58 at 4K. Both need a still plus audio. Pick resolution against cost.
Match the row to the files you have: audio-to, image-to, or video-to. Credits follow the inputs.
DreamActor maps a driving face onto a still. Kling Avatar Pro lipsyncs a track onto a still. Pick the file you actually have.
Run Kling Avatar Pro when the cut shows a mouth. Skip it for VO under b-roll, scene invent, or face transfer from a driving clip.
Wan 2.2 Speech to Video is 720p lipsync from a still and a track. It will not invent a presenter from a sentence.
Align captions to the audio, not the generated mouth. Recaption after every lipsync pass; the old track sits on the old take.
Pick a lipsync model by language, face fidelity and whether you need dubs, talking photos or avatar reads.
Time captions to the final audio after lipsync. Captioning a pre-lipsync scratch track guarantees drift.
Wan 2.7 T2V and I2V rows in catalog carry null audio. Plan TTS and lipsync as separate steps when speech is required.
The leftover /cost/dubbing-a-video-into-10-languages meter is job × languages. Which language you picked, how many speakers are in the source, and the lipsync per-second rate card do not multiply. A resubmission is a new job. Japanese and Spanish are the same charge.
Avatar X Text to Video wants a plate and a track. It is not how you invent the scene.
HeyGen-class lipsync is 2.5x credits in the app. Use it when we see the mouth. Voice-only dub when we do not.
The app's example rail is one clip through lipsync dubs. That is the localisation product, not three reshoots.
A talking avatar is one lipsync job. Text-to-video of a 'person speaking' still has no mouth contract and spends the wrong model.
Kling Avatar Pro wants a plate and a track. It is not how you invent the scene.
Talking-avatar examples in the app. Use them when the deliverable is a mouth saying the line, not a world.
LTX 2.3 Audio to Video wants a plate and a track. It is not how you invent the scene.
VEED Fabric 1.0 Text wants a plate and a track. It is not how you invent the scene.
VEED Lipsync wants a plate and a track. It is not how you invent the scene.
Wan 2.2 Speech to Video wants a plate and a track. It is not how you invent the scene.
generate_lipsync needs a face image and an audio file. Driving a rejected still, or a scratch read, spends a lipsync model on inputs you will replace.
Video models render at 24fps and Versely timelines default to 25. The three ways to conform, the arithmetic behind each, and the one that wrecks lip sync.
Generating picture from an audio track or scoring picture after it exists changes what you can still fix. A rule keyed to whichever element is locked.