Three input shapes exist and they behave differently. Audio plus a still photo generates the whole performance, including head movement it invented. Audio plus existing footage repaints only the mouth region and leaves your original performance intact. Text plus a face collapses the two steps, generating the voice and the mouth in one pass.
The quality ceiling is set by the face, not the audio. Frontal, evenly lit, mouth clearly visible, no hand crossing the jaw — those inputs sync cleanly. Profile angles, heavy shadow and fast head turns are where the mouth region starts to smear.
Lipsync is one component of dubbing, not a synonym for it. Dubbing also involves translation and a voice; lipsync is only the part that makes the mouth agree with whatever audio you ended up with.
In practice
- Audio quality drives sync accuracy — background music bleeding into the vocal track confuses the alignment.
- Clip length is bounded by the audio, so trim the take before generating rather than after.
- Video-input lipsync preserves the original body performance; photo-input lipsync invents one.
Lipsync models
Catalog entries whose whole job is matching a mouth to a soundtrack. 19 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| HeyGen Avatar V5 | HeyGen | Lipsync |
| LTX 2.3 Audio to Video | LTX | Lipsync |
| LTX 2 Audio to Video | LTX | Lipsync |
| Kling Avatar Pro | Kling | Lipsync |
| Sync React 1 | Sync | Lipsync |
| VEED Avatars | VEED | Lipsync |
| VEED Fabric 1.0 | VEED | Lipsync |
| VEED Fabric 1.0 Text | VEED | Lipsync |
Browse all 13 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Using a source face that moves too much. The mouth region has to be tracked frame by frame, and a fast turn or an occluding hand is where the artefacts appear.
Where you will run into it
- Lipsync a Photo or Video to Audio — One photo, one audio file, one talking clip.
- Add Multi-Speaker Dialogue to a Video — A whole cast, one API call.
- AI Lipsync Generator — Text, audio or video in. Pixel-perfect talking head out.
Related terms
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
Audio-to-video
Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.
Forced alignment
Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.
Character consistency
Character consistency is whether the same person, mascot or product still looks like itself across separate generations.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.