Guides

    Animate a portrait to speech after the still and the line are locked

    generate_lipsync needs a face image and an audio file. Driving a rejected still, or a scratch read, spends a lipsync model on inputs you will replace.

    Versely Team3 min read

    generate_lipsync animates a still face to an audio file. Image URL in, audio URL in, talking clip out. It is not a video-to-video mouth repair. If the still is a maybe, or the read is scratch, you spent a lipsync model on two inputs that will both move.

    Lipsync a photo or video to audio is the page — the "video" in the slug is the output, not the face input. The AI lipsync tool is the catalog door. Models you can actually pick include VEED Lipsync (4 credits), Sync Lipsync 2.0 (25), and VEED Fabric 1.0 Text (40). That spread is the whole argument for not tasting on a miss.

    Two locks, not one

    The portrait has to be the portrait: clear, front-facing, identity signed. Side profiles and busy occlusions degrade. Do not lipsync the rejected headshot "to see if the model is good." Test the model on the still you will actually use, or do not test it yet.

    The audio has to be the audio: words, duration, voice. A new script is a new performance. Lipsync does not make a bad still into a presenter. It makes a presenter from a still you already believe.

    If what you wanted was to retarget the mouth on an existing talking video to a new language, that is not this tool. That is dub_video with the HeyGen engine, and it has the same later-step rule.

    Do not buy the 40-credit row to audition a maybe

    Fabric at 40 versus VEED Lipsync at 4 is a 10× spread. Auditioning Fabric on a throwaway still is how a credit balance disappears with no presenter to ship. Pick the model for the job after the inputs are keepers. A 4-credit pass on a locked still and a locked read is information. A 40-credit pass on a maybe is a souvenir.

    UGC talking heads that already live in a room should often be generated as video, not still-plus-lipsync. UGC video is that path. Lipsync is for the case where the identity is a photo you control and the line is audio you control.

    Order

    1. Lock the still. Identity, crop, wardrobe.
    2. Lock the read. generate_speech or a real VO, duration that matches the intended clip.
    3. Run generate_lipsync with a model you chose on purpose.
    4. Caption and crop after the talking clip exists, because the mouth pass is now the picture.

    FAQ

    Can I feed it an existing talking video?

    Not as the face input. Image plus audio. Video mouth repair for a new language is HeyGen dubbing.

    Should I upscale the still first?

    Only if this still is the still. Upscaling a rejected headshot then lipsyncing it is two finishing meters on a miss. Image upscale after you know the portrait is the one you will drive.

    Does a cloned voice change the rule?

    No. Clone once from a clean sample you own. Drive keepers with that voice_id. Driving every experimental still with the clone still spends the lipsync model per clip.

    Why is this a later step if the photo is already done?

    Because the audio usually is not. Teams lipsync a temp line, then rewrite the offer, then pay the model again. Lock the sentence.