How to pick a lipsync model
Pick a lipsync model by language, face fidelity and whether you need dubs, talking photos or avatar reads.
Every guide, comparison and workflow we’ve published on AI Lipsync.
39 articles — page 1 of 2
Pick a lipsync model by language, face fidelity and whether you need dubs, talking photos or avatar reads.
Time captions to the final audio after lipsync. Captioning a pre-lipsync scratch track guarantees drift.
Wan 2.7 T2V and I2V rows in catalog carry null audio. Plan TTS and lipsync as separate steps when speech is required.
veed-avatars-lipsync is a premade-avatars row at 30 credits and 4K — you pick a stock presenter; it is not photo-in Avatar 4 at 50 credits.
The leftover /cost/dubbing-a-video-into-10-languages meter is job × languages. Which language you picked, how many speakers are in the source, and the lipsync per-second rate card do not multiply. A resubmission is a new job. Japanese and Spanish are the same charge.
Avatar X Text to Video wants a plate and a track. It is not how you invent the scene.
HeyGen-class lipsync is 2.5x credits in the app. Use it when we see the mouth. Voice-only dub when we do not.
The app's example rail is one clip through lipsync dubs. That is the localisation product, not three reshoots.
Gemini 3.1 Flash TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
A talking avatar is one lipsync job. Text-to-video of a 'person speaking' still has no mouth contract and spends the wrong model.
HeyGen Avatar V3 wants a plate and a track. It is not how you invent the scene.
HeyGen Avatar V5 wants a plate and a track. It is not how you invent the scene.
HeyGen Image to Video wants a plate and a track. It is not how you invent the scene.
Kling Avatar Pro wants a plate and a track. It is not how you invent the scene.
Kling Lipsync wants a plate and a track. It is not how you invent the scene.
Talking-avatar examples in the app. Use them when the deliverable is a mouth saying the line, not a world.
LTX 2.3 Audio to Video wants a plate and a track. It is not how you invent the scene.
LTX 2 Audio to Video wants a plate and a track. It is not how you invent the scene.
Sync React 1 wants a plate and a track. It is not how you invent the scene.
VEED Avatars wants a plate and a track. It is not how you invent the scene.
VEED Fabric 1.0 Text wants a plate and a track. It is not how you invent the scene.
VEED Fabric 1.0 wants a plate and a track. It is not how you invent the scene.
VEED Lipsync wants a plate and a track. It is not how you invent the scene.
Wan 2.2 Speech to Video wants a plate and a track. It is not how you invent the scene.