Audio, Voice & Dubbing · Versely AI

    Lipsync a Photo to Audio — Make a Still Image Talk

    One photo, one audio file, one talking clip.

    generate_lipsync animates a face IMAGE to speak provided audio — it's the tool behind AI presenters, talking-avatar UGC, and animated portraits. It needs an image_url and an audio_url; it doesn't take an existing video as input.

    If what you actually want is to re-sync the mouth movement on an EXISTING video to a new language's audio, that's a different job — see dub-video, whose 'heygen' engine does lip-synced translation on real footage rather than a still photo.

    Powered by

    generate_lipsync

    Generate a lipsync video — animate a face image to speak with provided audio. Use when the user wants to make a person in an image speak, create a talking avatar, or lip-sync audio to a face.

    What to tell the agent

    Versely's agent maps this plain-English request directly onto generate_lipsync. You don't need to know the parameter names — just describe what you want.

    Animate this portrait photo to speak this audio file, natural expression, medium talking style.

    How it works

    1. 1. Pick a face image

      A single, clear, front-facing portrait works best.

    2. 2. Supply the audio

      Either an existing audio_url (a recording or a prior generate_speech output), or generate one first if you're starting from a script.

    3. 3. Pick a model and tune delivery

      emotion, expression, talking_style, resolution and sync_mode are all controllable depending on which lipsync model you choose.

    4. 4. Generate

      Versely aligns phonemes to mouth shapes and returns a finished talking clip.

    Models & real cost

    Cost depends on the model you pick. Real options include VEED Lipsync (4 credits, video-to-lipsync), Sync Lipsync 2.0 (25 credits, video-to-lipsync), and VEED Fabric 1.0 Text (40 credits, text- and image-to-lipsync) — despite the category names, all three are selectable from generate_lipsync's model parameter for animating a face to speak.

    The formula behind that number: What does a 30-second AI talking head cost? bills a flat rate for every second of output, so length is the only lever.

    Limits & things to know

    • generate_lipsync takes an image, not an existing video, as the face input — for re-syncing an already-shot video's mouth to new audio, use dub_video's 'heygen' engine instead.
    • Works best on clear, front-facing portraits; side profiles and heavily obscured faces degrade quality.
    • Credit cost varies significantly by model — VEED Lipsync at 4 credits vs VEED Fabric 1.0 Text at 40 credits is a 10x spread, so pick deliberately.

    Who uses this

    • AI presenter / faceless-channel avatars
    • Turning a product photo's model into a talking spokesperson
    • Personalized outreach videos at scale
    • Historical figure or illustrated-character explainer videos

    Frequently asked questions

    Can I lipsync an existing video, not just a photo?+

    generate_lipsync's input is an image, not a video. To re-sync an existing video's mouth movement to new audio, use dub_video with engine 'heygen' instead — that's the video-native lip-sync path.

    What models power lipsync in Versely?+

    Real options include VEED Lipsync (4 credits), Sync Lipsync 2.0 (25 credits), and VEED Fabric 1.0 Text (40 credits) — pricing and quality vary, so check each before picking.

    Does the face need to be a real photo?+

    Most front-facing portraits work, including some illustrations and stylized characters. Side profiles and heavily obscured faces typically degrade the result.

    Can I control the emotion of the delivery?+

    Yes — emotion, expression and talking_style parameters are available, though exact support depends on which model you select.

    Related Versely tools

    Related editing jobs

    Lipsync a Photo or Video to Audio inside Versely

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.