Wan · talking-head & lipsync video

    Wan 2.2 Speech to Video Prompting Guide

    What this model actually wants — from its schema, not from vibes.

    Wan 2.2 Speech to Video is Wan's talking-head & lipsync video model on Versely. This page is its structured prompting reference: the 5 parameters its schema actually exposes, the talking-video technique that applies to it, copy-ready templates.

    Everything here is grounded in the same sources Versely's agent reads — the model's input schema, the wan family rule its prompt enhancer applies. Where a line is general craft advice rather than a documented fact about Wan 2.2 Speech to Video, the page says so.

    What Wan 2.2 Speech to Video wants

    The exact input surface, from the same schema the Versely agent fetches with get_model_input_schema before every generation.

    ParameterWhat it doesValues
    image_urlreqReference portrait/character imagestring
    audio_urlreqDriving speech audiostring
    promptOptional scene/motion promptstring
    resolutionResolution480p · 720pdefault: 720p
    durationSeconds (defaults to audio length)integer
    • Schema skeleton — verify before strict mode

    The rule Versely's enhancer applies

    Versely's prompt enhancer carries a per-family rule for wan models, applied automatically whenever it rewrites a prompt for Wan 2.2 Speech to Video. Verbatim:

    Wan models work well with straightforward scene descriptions and style keywords.

    Technique that applies here

    Lipsync/avatar/audio-driven: text or audio input, delivery and framing controls

    • Three different input shapes share this family, and figuring out which one your model uses tells you what 'the prompt' even means: text-is-the-script (HeyGen Avatar V3, VEED Avatars — your text becomes the spoken words), audio-drives-it (Kling Avatar Pro, LTX 2.3 Audio to Video, Wan 2.2 Speech to Video — the words come from your uploaded audio_url; any prompt text only shapes the visual scene/motion around it), or pure re-timing with no prompt field at all (Sync Lipsync 2.0, VEED Lipsync, VEED Fabric 1.0, Sync React 1).
    • Kling Avatar Pro's input.prompt is required (max 5000 chars), but the schema describes it as a 'Motion/scene description,' not dialogue — writing your script into it does nothing, since input.audio_url is what the avatar actually lip-syncs to. Wan 2.2 Speech to Video has the same audio-drives-it shape but makes its equivalent prompt field optional.
    • The enhancer's family tips still apply in lipsync mode: Kling Avatar Pro gets the kling MODEL_TIPS text ('clear scene descriptions including camera movement..., subject action, and environment details') and Wan 2.2 Speech to Video gets the wan text ('straightforward scene descriptions and style keywords') — both aimed at the non-verbal scene/motion half of the prompt, since the dialogue itself is audio-driven on both.
    • General technique: when a model does take a scene/motion prompt alongside driving audio (Kling, LTX, Wan), describe body language and framing that could plausibly match any speech — you're describing the performance, not the words, and the two need to stay decoupled or the visual will fight the audio track.

    Copy-ready templates

    Replace the bracketed slots; each template says when it's the right shape.

    Template 1
    voice.prompt: "[SCRIPT — what the avatar says, written as natural spoken sentences]" · character.avatar: [AVATAR ID], voice.voice: [VOICE ID], resolution: [720p/1080p], output_language: [LANGUAGE CODE if dubbing]

    Use when: Scripting a HeyGen Avatar V3 talking head from text — no separate audio file needed, HeyGen generates the voice.

    Template 2
    [AVATAR] [GESTURE OR ACTION] while [FRAMING: close-up/medium/wide], in [SETTING], [LIGHTING NOTE]

    Use when: Writing the scene/motion prompt for Kling Avatar Pro, LTX 2.3 Audio to Video, or Wan 2.2 Speech to Video, where the spoken words come from your uploaded audio_url, not this text.

    How the Versely agent does this automatically

    You can use this page by hand, or let the agent apply the same knowledge. Four real mechanisms — no more, no less:

    • get_model_input_schema — before generating, the agent looks up Wan 2.2 Speech to Video's exact input fields, required fields, allowed values, defaults, and min/max bounds. The parameter table above is that same surface.
    • The prompt enhancer's family rules — 12 per-family rewrite rules, including the wan rule quoted on this page, shape how a rough prompt gets rewritten.
    • The per-provider speech guide — for TTS scripts, the agent follows a provider-specific tag scheme — not relevant to this model, but it's why voiceover scripts come out marked up correctly.
    • expand_movie_scene — in movie flows, brief scene ideas are rewritten into detailed cinematic descriptions before generation.

    Mistakes that waste generations

    • Typing your script into Kling Avatar Pro's or Wan 2.2 Speech to Video's prompt field and expecting the avatar to say it — both prompts are scene/motion descriptions; the spoken words come only from audio_url.
    • Writing a phrase like 'a little sad but trying to smile' into Sync React 1's emotion field — the schema requires exactly one word from a fixed 6-value enum (happy/angry/sad/neutral/disgusted/surprised); anything else is invalid.
    • Assuming VEED Avatars supports a separate voice or language selector the way HeyGen Avatar V3 does — its schema is explicitly limited to avatar_id and text; voice_id, language, and aspect_ratio are documented as not valid fields.

    Frequently asked questions

    Does Wan 2.2 Speech to Video support negative prompts?+

    No — Wan 2.2 Speech to Video's published schema has no negative_prompt parameter. Exclusions have to be phrased positively inside the main prompt, or dropped.

    How does the Versely agent know Wan 2.2 Speech to Video's parameters?+

    Before generating, the agent calls its get_model_input_schema tool, which looks up the exact input fields, required fields, allowed values, defaults, and min/max bounds for the model; separately, the prompt enhancer applies the wan family rule quoted on this page to the prompt text itself. Nothing on this page is guessed — it is the same schema surface those tools read.

    Related prompting guides

    Generate with Wan 2.2 Speech to Video

    Wan 2.2 Speech to Video is live in Versely — paste a template above, or just describe what you want and let the agent map it onto the schema for you.