Guides

    generate_lipsync needs a face and audio, not a text-to-video prompt

    A talking avatar is one lipsync job. Text-to-video of a 'person speaking' still has no mouth contract and spends the wrong model.

    Versely Team4 min read

    generate_lipsync needs a face and audio. It is not text-to-video of a person who happens to be talking. It is not a voiceover. It is not a dub. Prompt a cinematic talking close-up on a video model and you get a mouth that approximates speech. You do not get lipsync. You get a clip you cannot caption to a script you never recorded.

    Make a talking avatar video is the named agent job. Photo plus audio in; a talking video out. The tool animates the face to speak the provided audio, matching mouth movement to the words. Required: a model, an image_url, an audio_url. Optional emotion, expression, talking style, resolution, sync mode. None of those optionals replace the two inputs.

    You can ask for a HeyGen digital-twin avatar instead of your own still; listing avatars and voices is supporting work. The job is still lipsync. It is still not "invent a presenter from a sentence."

    What the job is

    A still (or listed avatar) that already looks like the person you want on camera. Audio that already says the line — generated speech, a clone, a recording. Then a lipsync model priced per model, spanning HeyGen avatars and other talking-head rows, shown before you confirm.

    The AI lipsync tool is the same contract as a launcher. This capability is the agent name for it.

    If you do not have audio yet, make audio first: write and generate a voiceover, or clone a voice and then speak the script. Do not ask lipsync to write the line. It has no text field that replaces audio_url.

    What it is not

    • Text-to-video "talking presenter." generate_videos will perform a vibe. It will not lock phonemes. Using it because it is cheaper-looking than a lipsync row is how you buy a silent chew.
    • Image-to-video of a portrait. Turn a photo into a video moves the still. It does not speak. "Smile and say the tagline" in that prompt is fan fiction.
    • Dubbing an existing video. Dub a video into another language starts from hosted video, not a still plus a new VO. Different engines, different limits, different job.
    • Sign, sketch, or motion-control dance. Mouth-to-audio is the objective. Pointing this stack at a job that is not speech does not change the objective.

    Each miss still spends a lipsync (or video) credit. The file you wanted remains unmade.

    The test

    Do you already have a face you approve and audio you approve, and the only remaining work is the mouth?

    Yes: this job. Attach both. Name the model if you care which engine. Confirm.

    If you have only a script, make speech. If you have only a vibe, make a still, then speech, then this. Skip the order and you will pay the last step twice.

    FAQ

    Can I type the script straight into the lipsync capability?

    Not as a substitute for audio_url. Generate or record the VO, then lipsync. Combining "write the line and also talk" in one sentence is how the agent has to split a job you pretended was one.

    Is a video model with native audio the same thing?

    Native audio on a scene model is ambience and sometimes speech-shaped sound. It is not a locked read of your copy on your face. If the mouth is the product, stay on generate_lipsync.

    Do I need HeyGen?

    No. HeyGen is one family on the row. The job is still: face, audio, lipsync model. Pick the catalog line that matches the still you have.

    Why not generate a talking clip and caption it to fake the read?

    You can upload that. You cannot defend the words. Captions on invented mouth flaps are a different finishing job on a file that never had a script.