What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Give the agent a face image and the audio it should speak — or ask for a HeyGen digital-twin avatar and it can list the available ones first.
- 2
It animates the face to speak the provided audio, matching mouth movement to the words.
- 3
You get back a lipsynced talking video, ready for UGC-style content or an explainer.
What it needs from you
- •Which AI model to use
- •An image to start from
- •An audio clip to work from
What comes back
A video of the face speaking the given audio, lips synced to the words.
What it costs
Priced per model — HeyGen avatars and lipsync models span a wide range on Versely.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
generate_lipsynclist_heygen_avatarslist_heygen_voiceslist_avatarsGenerate a lipsync video — animate a face image to speak with provided audio. Use when the user wants to make a person in an image speak, create a talking avatar, or lip-sync audio to a face.
The full tool behind it
See it done in a real workflow
Morning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
TrendingNYC Street Interview
Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
TrendingPrimo Protein vs Other Brand
Pixar-style 3D comparison ad — the confident PRIMO PROTEIN pouch faces off against the tired old rival brand's tub across 11 talking clips: pasture vs dusty pantry, herb garden vs toxic lab, clean American lab vs grimy factory, chocolate-milkshake CTA vs "wet sand." ~40s vertical reel with native character voices.
Or start from a one-tap template
Frequently asked questions
What do I actually say to the agent to make a talking avatar video?+
Just describe it in plain English — for example: "[Attach a face photo and audio] Make this person say this" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: Which AI model to use; An image to start from; An audio clip to work from. Anything else it needs, it asks for before running.
What do I get back?+
A video of the face speaking the given audio, lips synced to the words.
Does this cost credits?+
Priced per model — HeyGen avatars and lipsync models span a wide range on Versely.
Is this one single action, or several?+
Behind the scenes the agent may call more than one tool to pull this off — but you only ever describe the outcome you want in one message.
You can also just ask for
Ask your Versely agent to make a talking avatar video
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.