Generation modes

    Lipsync meaning

    Also called Lip sync, Talking head generation.

    Lipsync meaning: driving a face mouth from an audio track so speech reads as spoken, not dubbed over the top.

    Three input shapes exist and they behave differently. Audio plus a still photo generates the whole performance, including head movement it invented. Audio plus existing footage repaints only the mouth region and leaves your original performance intact. Text plus a face collapses the two steps, generating the voice and the mouth in one pass.

    The quality ceiling is set by the face, not the audio. Frontal, evenly lit, mouth clearly visible, no hand crossing the jaw — those inputs sync cleanly. Profile angles, heavy shadow and fast head turns are where the mouth region starts to smear.

    Lipsync is one component of dubbing, not a synonym for it. Dubbing also involves translation and a voice; lipsync is only the part that makes the mouth agree with whatever audio you ended up with.

    In practice

    • Audio quality drives sync accuracy — background music bleeding into the vocal track confuses the alignment.
    • Clip length is bounded by the audio, so trim the take before generating rather than after.
    • Video-input lipsync preserves the original body performance; photo-input lipsync invents one.

    Lipsync models

    Catalog entries whose whole job is matching a mouth to a soundtrack. 23 of the 331 models in the Versely catalog qualify.

    ModelProviderType
    LTX 2.5 Audio to Video ProLTXLipsync
    Avatar X Text to VideoMirageLipsync
    HeyGen Avatar V5HeyGenLipsync
    LTX 2.3 Audio to VideoLTXLipsync
    LTX 2 Audio to VideoLTXLipsync
    Kling Avatar ProKlingLipsync
    Kling LipsyncKlingLipsync
    Sync React 1SyncLipsync

    Browse all 16 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Using a source face that moves too much. The mouth region has to be tracked frame by frame, and a fast turn or an occluding hand is where the artefacts appear.

    Go deeper

    Inworld TTS-2 for dubs, then a lipsync model

    Inworld TTS-2 is the high-volume dub voice, priced under ElevenLabs v3. ElevenLabs v3 is the premium read. Sync.so or Hedra is the lipsync model after that.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.

    All terms →