HeyGen · talking-head & lipsync video

    HeyGen Avatar V3 Prompting Guide

    What this model actually wants — from its schema, not from vibes.

    HeyGen Avatar V3 is HeyGen's talking-head & lipsync video model on Versely. This page is its structured prompting reference: the 3 parameters its schema actually exposes, the talking-video technique that applies to it, copy-ready templates.

    Everything here is grounded in the same sources Versely's agent reads — the model's input schema. Where a line is general craft advice rather than a documented fact about HeyGen Avatar V3, the page says so.

    What HeyGen Avatar V3 wants

    The exact input surface, from the same schema the Versely agent fetches with get_model_input_schema before every generation.

    ParameterWhat it doesValues
    characterreqAvatar referenceobject
    voicereqVoice + scriptobject
    resolutionResolution720p · 1080pdefault: 720p
    • Nested character/voice objects per controller

    Technique that applies here

    Lipsync/avatar/audio-driven: text or audio input, delivery and framing controls

    • Three different input shapes share this family, and figuring out which one your model uses tells you what 'the prompt' even means: text-is-the-script (HeyGen Avatar V3, VEED Avatars — your text becomes the spoken words), audio-drives-it (Kling Avatar Pro, LTX 2.3 Audio to Video, Wan 2.2 Speech to Video — the words come from your uploaded audio_url; any prompt text only shapes the visual scene/motion around it), or pure re-timing with no prompt field at all (Sync Lipsync 2.0, VEED Lipsync, VEED Fabric 1.0, Sync React 1).
    • VEED Avatars' schema is narrower than sibling avatar products: only avatar_id and text are valid fields, per the schema's own constraint note — voice_id, language, and aspect_ratio are explicitly documented as not supported, even though HeyGen Avatar V3 does expose a separate voice ID and resolution control.
    • HeyGen Avatar V3 nests its fields two levels deep — character.avatar (+ optional avatar_style) for who's on screen, voice.prompt (the actual script) and voice.voice (voice ID) for what's said and how, plus a top-level output_language to dub the same script into another language without touching the avatar or voice IDs.

    Copy-ready templates

    Replace the bracketed slots; each template says when it's the right shape.

    Template 1
    voice.prompt: "[SCRIPT — what the avatar says, written as natural spoken sentences]" · character.avatar: [AVATAR ID], voice.voice: [VOICE ID], resolution: [720p/1080p], output_language: [LANGUAGE CODE if dubbing]

    Use when: Scripting a HeyGen Avatar V3 talking head from text — no separate audio file needed, HeyGen generates the voice.

    Template 2
    text: "[SCRIPT — what the avatar says]" · avatar_id: [PRESET AVATAR ID]

    Use when: Scripting VEED Avatars — remember only avatar_id and text exist here; there's no separate voice, language, or aspect_ratio field to set.

    How the Versely agent does this automatically

    You can use this page by hand, or let the agent apply the same knowledge. Four real mechanisms — no more, no less:

    • get_model_input_schema — before generating, the agent looks up HeyGen Avatar V3's exact input fields, required fields, allowed values, defaults, and min/max bounds. The parameter table above is that same surface.
    • The prompt enhancer's family rules — 12 per-family rewrite rules (this model's family isn't one of the 12, so only general enhancement applies) shape how a rough prompt gets rewritten.
    • The per-provider speech guide — for TTS scripts, the agent follows a provider-specific tag scheme — not relevant to this model, but it's why voiceover scripts come out marked up correctly.
    • expand_movie_scene — in movie flows, brief scene ideas are rewritten into detailed cinematic descriptions before generation.

    Mistakes that waste generations

    • Typing your script into Kling Avatar Pro's or Wan 2.2 Speech to Video's prompt field and expecting the avatar to say it — both prompts are scene/motion descriptions; the spoken words come only from audio_url.
    • Writing a phrase like 'a little sad but trying to smile' into Sync React 1's emotion field — the schema requires exactly one word from a fixed 6-value enum (happy/angry/sad/neutral/disgusted/surprised); anything else is invalid.
    • Assuming VEED Avatars supports a separate voice or language selector the way HeyGen Avatar V3 does — its schema is explicitly limited to avatar_id and text; voice_id, language, and aspect_ratio are documented as not valid fields.

    This guide also covers

    These siblings share HeyGen Avatar V3's prompting-relevant input surface, so their prompting URLs resolve here — tier and pricing differences live on their own model pages:

    Frequently asked questions

    Does HeyGen Avatar V3 support negative prompts?+

    No — HeyGen Avatar V3's published schema has no negative_prompt parameter. Exclusions have to be phrased positively inside the main prompt, or dropped.

    How does the Versely agent know HeyGen Avatar V3's parameters?+

    Before generating, the agent calls its get_model_input_schema tool, which looks up the exact input fields, required fields, allowed values, defaults, and min/max bounds for the model. Nothing on this page is guessed — it is the same schema surface those tools read.

    Does this guide also cover HeyGen Avatar V5?+

    Yes. HeyGen Avatar V5 share the same prompting-relevant input surface as HeyGen Avatar V3, so their prompting URLs redirect here instead of duplicating this page. Tier and pricing differences live on each model's own /models page.

    Related prompting guides

    Generate with HeyGen Avatar V3

    HeyGen Avatar V3 is live in Versely — paste a template above, or just describe what you want and let the agent map it onto the schema for you.