Three input shapes exist and they behave differently. Audio plus a still photo generates the whole performance, including head movement it invented. Audio plus existing footage repaints only the mouth region and leaves your original performance intact. Text plus a face collapses the two steps, generating the voice and the mouth in one pass.
The quality ceiling is set by the face, not the audio. Frontal, evenly lit, mouth clearly visible, no hand crossing the jaw — those inputs sync cleanly. Profile angles, heavy shadow and fast head turns are where the mouth region starts to smear.
Lipsync is one component of dubbing, not a synonym for it. Dubbing also involves translation and a voice; lipsync is only the part that makes the mouth agree with whatever audio you ended up with.
In practice
- Audio quality drives sync accuracy — background music bleeding into the vocal track confuses the alignment.
- Clip length is bounded by the audio, so trim the take before generating rather than after.
- Video-input lipsync preserves the original body performance; photo-input lipsync invents one.
Lipsync models
Catalog entries whose whole job is matching a mouth to a soundtrack. 23 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| LTX 2.5 Audio to Video Pro | LTX | Lipsync |
| Avatar X Text to Video | Mirage | Lipsync |
| HeyGen Avatar V5 | HeyGen | Lipsync |
| LTX 2.3 Audio to Video | LTX | Lipsync |
| LTX 2 Audio to Video | LTX | Lipsync |
| Kling Avatar Pro | Kling | Lipsync |
| Kling Lipsync | Kling | Lipsync |
| Sync React 1 | Sync | Lipsync |
Browse all 16 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Using a source face that moves too much. The mouth region has to be tracked frame by frame, and a fast turn or an occluding hand is where the artefacts appear.
Go deeper
Inworld TTS-2 for dubs, then a lipsync model
Inworld TTS-2 is the high-volume dub voice, priced under ElevenLabs v3. ElevenLabs v3 is the premium read. Sync.so or Hedra is the lipsync model after that.
Where you will run into it
- Lipsync a Photo or Video to Audio — One photo, one audio file, one talking clip.
- Add Multi-Speaker Dialogue to a Video — A whole cast, one API call.
- AI Lipsync Generator — Text, audio or video in. Talking head out.
Related terms
AI dubbing
AI dubbing meaning: replacing spoken audio with another language, usually keeping the speaker voice and optionally re-syncing lips.
Audio-to-video
Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.
Forced alignment
Forced alignment meaning: matching a known transcript to its audio to find exactly when each word was spoken.
Character consistency
Character consistency meaning: whether the same person, mascot, or product still looks like itself across separate generations.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.