generate_lipsync animates a face IMAGE to speak provided audio — it's the tool behind AI presenters, talking-avatar UGC, and animated portraits. It needs an image_url and an audio_url; it doesn't take an existing video as input.
If what you actually want is to re-sync the mouth movement on an EXISTING video to a new language's audio, that's a different job — see dub-video, whose 'heygen' engine does lip-synced translation on real footage rather than a still photo.
Powered by
Generate a lipsync video — animate a face image to speak with provided audio. Use when the user wants to make a person in an image speak, create a talking avatar, or lip-sync audio to a face.
What to tell the agent
Versely's agent maps this plain-English request directly onto generate_lipsync. You don't need to know the parameter names — just describe what you want.
“Animate this portrait photo to speak this audio file, natural expression, medium talking style.”
How it works
1. Pick a face image
A single, clear, front-facing portrait works best.
2. Supply the audio
Either an existing audio_url (a recording or a prior generate_speech output), or generate one first if you're starting from a script.
3. Pick a model and tune delivery
emotion, expression, talking_style, resolution and sync_mode are all controllable depending on which lipsync model you choose.
4. Generate
Versely aligns phonemes to mouth shapes and returns a finished talking clip.
Models & real cost
Cost depends on the model you pick. Real options include VEED Lipsync (4 credits, video-to-lipsync), Sync Lipsync 2.0 (25 credits, video-to-lipsync), and VEED Fabric 1.0 Text (40 credits, text- and image-to-lipsync) — despite the category names, all three are selectable from generate_lipsync's model parameter for animating a face to speak.
The formula behind that number: What does a 30-second AI talking head cost? — bills a flat rate for every second of output, so length is the only lever.
Limits & things to know
- generate_lipsync takes an image, not an existing video, as the face input — for re-syncing an already-shot video's mouth to new audio, use dub_video's 'heygen' engine instead.
- Works best on clear, front-facing portraits; side profiles and heavily obscured faces degrade quality.
- Credit cost varies significantly by model — VEED Lipsync at 4 credits vs VEED Fabric 1.0 Text at 40 credits is a 10x spread, so pick deliberately.
Who uses this
- AI presenter / faceless-channel avatars
- Turning a product photo's model into a talking spokesperson
- Personalized outreach videos at scale
- Historical figure or illustrated-character explainer videos
Frequently asked questions
Can I lipsync an existing video, not just a photo?+
generate_lipsync's input is an image, not a video. To re-sync an existing video's mouth movement to new audio, use dub_video with engine 'heygen' instead — that's the video-native lip-sync path.
What models power lipsync in Versely?+
Real options include VEED Lipsync (4 credits), Sync Lipsync 2.0 (25 credits), and VEED Fabric 1.0 Text (40 credits) — pricing and quality vary, so check each before picking.
Does the face need to be a real photo?+
Most front-facing portraits work, including some illustrations and stylized characters. Side profiles and heavily obscured faces typically degrade the result.
Can I control the emotion of the delivery?+
Yes — emotion, expression and talking_style parameters are available, though exact support depends on which model you select.
Related Versely tools
Related editing jobs
Lipsync a Photo or Video to Audio inside Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.