generate_lipsync needs a face and audio, not a text-to-video prompt
A talking avatar is one lipsync job. Text-to-video of a 'person speaking' still has no mouth contract and spends the wrong model.
generate_lipsync needs a face and audio. It is not text-to-video of a person who happens to be talking. It is not a voiceover. It is not a dub. Prompt a cinematic talking close-up on a video model and you get a mouth that approximates speech. You do not get lipsync. You get a clip you cannot caption to a script you never recorded.
Make a talking avatar video is the named agent job. Photo plus audio in; a talking video out. The tool animates the face to speak the provided audio, matching mouth movement to the words. Required: a model, an image_url, an audio_url. Optional emotion, expression, talking style, resolution, sync mode. None of those optionals replace the two inputs.
You can ask for a HeyGen digital-twin avatar instead of your own still; listing avatars and voices is supporting work. The job is still lipsync. It is still not "invent a presenter from a sentence."
What the job is
A still (or listed avatar) that already looks like the person you want on camera. Audio that already says the line — generated speech, a clone, a recording. Then a lipsync model priced per model, spanning HeyGen avatars and other talking-head rows, shown before you confirm.
The AI lipsync tool is the same contract as a launcher. This capability is the agent name for it.
If you do not have audio yet, make audio first: write and generate a voiceover, or clone a voice and then speak the script. Do not ask lipsync to write the line. It has no text field that replaces audio_url.
What it is not
- Text-to-video "talking presenter."
generate_videoswill perform a vibe. It will not lock phonemes. Using it because it is cheaper-looking than a lipsync row is how you buy a silent chew. - Image-to-video of a portrait. Turn a photo into a video moves the still. It does not speak. "Smile and say the tagline" in that prompt is fan fiction.
- Dubbing an existing video. Dub a video into another language starts from hosted video, not a still plus a new VO. Different engines, different limits, different job.
- Sign, sketch, or motion-control dance. Mouth-to-audio is the objective. Pointing this stack at a job that is not speech does not change the objective.
Each miss still spends a lipsync (or video) credit. The file you wanted remains unmade.
The test
Do you already have a face you approve and audio you approve, and the only remaining work is the mouth?
Yes: this job. Attach both. Name the model if you care which engine. Confirm.
If you have only a script, make speech. If you have only a vibe, make a still, then speech, then this. Skip the order and you will pay the last step twice.
FAQ
Can I type the script straight into the lipsync capability?
Not as a substitute for audio_url. Generate or record the VO, then lipsync. Combining "write the line and also talk" in one sentence is how the agent has to split a job you pretended was one.
Is a video model with native audio the same thing?
Native audio on a scene model is ambience and sometimes speech-shaped sound. It is not a locked read of your copy on your face. If the mouth is the product, stay on generate_lipsync.
Do I need HeyGen?
No. HeyGen is one family on the row. The job is still: face, audio, lipsync model. Pick the catalog line that matches the still you have.
Why not generate a talking clip and caption it to fake the read?
You can upload that. You cannot defend the words. Captions on invented mouth flaps are a different finishing job on a file that never had a script.