Sync · talking-head & lipsync video

    Sync React 1 Prompting Guide

    What this model actually wants — from its schema, not from vibes.

    Sync React 1 is Sync's talking-head & lipsync video model on Versely. This page is its structured prompting reference: the 3 parameters its schema actually exposes, the talking-video technique that applies to it, copy-ready templates.

    Everything here is grounded in the same sources Versely's agent reads — the model's input schema. Where a line is general craft advice rather than a documented fact about Sync React 1, the page says so.

    What Sync React 1 wants

    The exact input surface, from the same schema the Versely agent fetches with get_model_input_schema before every generation.

    ParameterWhat it doesValues
    video_urlreqSource video (≤15s)string
    audio_urlreqDriving audio (≤15s)string
    emotionreqReaction emotion (REQUIRED per docs; single-word only)happy · angry · sad · neutral · disgusted · surprised
    • Verified against fal.ai/models/fal-ai/sync-lipsync/react-1 — model_mode enum is lips/face/head (NOT standard/pro); emotion required

    Technique that applies here

    Lipsync/avatar/audio-driven: text or audio input, delivery and framing controls

    • Three different input shapes share this family, and figuring out which one your model uses tells you what 'the prompt' even means: text-is-the-script (HeyGen Avatar V3, VEED Avatars — your text becomes the spoken words), audio-drives-it (Kling Avatar Pro, LTX 2.3 Audio to Video, Wan 2.2 Speech to Video — the words come from your uploaded audio_url; any prompt text only shapes the visual scene/motion around it), or pure re-timing with no prompt field at all (Sync Lipsync 2.0, VEED Lipsync, VEED Fabric 1.0, Sync React 1).
    • Kling Avatar Pro's input.prompt is required (max 5000 chars), but the schema describes it as a 'Motion/scene description,' not dialogue — writing your script into it does nothing, since input.audio_url is what the avatar actually lip-syncs to. Wan 2.2 Speech to Video has the same audio-drives-it shape but makes its equivalent prompt field optional.
    • Sync React 1 replaces a text prompt with enum dials: emotion (required, one word only, from happy/angry/sad/neutral/disgusted/surprised), model_mode (lips/face/head — how much of the frame reacts, default face), temperature (0–1 — how exaggerated the reaction is, default 0.5), and lipsync_mode (default bounce). Sync Lipsync 2.0 exposes that same length-mismatch choice as sync_mode (cut_off/loop/bounce/silence/remap), defaulting to cut_off instead.

    Copy-ready templates

    Replace the bracketed slots; each template says when it's the right shape.

    Template 1
    voice.prompt: "[SCRIPT — what the avatar says, written as natural spoken sentences]" · character.avatar: [AVATAR ID], voice.voice: [VOICE ID], resolution: [720p/1080p], output_language: [LANGUAGE CODE if dubbing]

    Use when: Scripting a HeyGen Avatar V3 talking head from text — no separate audio file needed, HeyGen generates the voice.

    Template 2
    emotion: [happy/angry/sad/neutral/disgusted/surprised] · model_mode: [lips/face/head] · temperature: [0.0–1.0] · lipsync_mode: [cut_off/loop/bounce/silence/remap]

    Use when: Dialing in a reaction clip on Sync React 1 — there's no free-text prompt field, only these four enum controls.

    How the Versely agent does this automatically

    You can use this page by hand, or let the agent apply the same knowledge. Four real mechanisms — no more, no less:

    • get_model_input_schema — before generating, the agent looks up Sync React 1's exact input fields, required fields, allowed values, defaults, and min/max bounds. The parameter table above is that same surface.
    • The prompt enhancer's family rules — 12 per-family rewrite rules (this model's family isn't one of the 12, so only general enhancement applies) shape how a rough prompt gets rewritten.
    • The per-provider speech guide — for TTS scripts, the agent follows a provider-specific tag scheme — not relevant to this model, but it's why voiceover scripts come out marked up correctly.
    • expand_movie_scene — in movie flows, brief scene ideas are rewritten into detailed cinematic descriptions before generation.

    Mistakes that waste generations

    • Typing your script into Kling Avatar Pro's or Wan 2.2 Speech to Video's prompt field and expecting the avatar to say it — both prompts are scene/motion descriptions; the spoken words come only from audio_url.
    • Writing a phrase like 'a little sad but trying to smile' into Sync React 1's emotion field — the schema requires exactly one word from a fixed 6-value enum (happy/angry/sad/neutral/disgusted/surprised); anything else is invalid.
    • Assuming VEED Avatars supports a separate voice or language selector the way HeyGen Avatar V3 does — its schema is explicitly limited to avatar_id and text; voice_id, language, and aspect_ratio are documented as not valid fields.

    Frequently asked questions

    Does Sync React 1 support negative prompts?+

    No — Sync React 1's published schema has no negative_prompt parameter. Exclusions have to be phrased positively inside the main prompt, or dropped.

    How does the Versely agent know Sync React 1's parameters?+

    Before generating, the agent calls its get_model_input_schema tool, which looks up the exact input fields, required fields, allowed values, defaults, and min/max bounds for the model. Nothing on this page is guessed — it is the same schema surface those tools read.

    Related prompting guides

    Generate with Sync React 1

    Sync React 1 is live in Versely — paste a template above, or just describe what you want and let the agent map it onto the schema for you.