Guides

    Wan 2.7 Prompting: References, Frames, and Voice

    Wan 2.7 prompting guide across all three modes: text-to-video scenes, image-to-video frame direction, and reference-to-video with voice-driven speech.

    Versely Team7 min read

    Wan 2.7 isn't one model to prompt — it's three, and writing the same kind of prompt for all of them is the most common way creators underuse it. The text-to-video mode wants a full scene described from nothing. The image-to-video mode wants direction for a frame that already exists. And the reference-to-video mode — the one that's quietly become a favorite for talking, persona-driven content — wants you to stop describing your subject entirely and spend the prompt on performance. This Wan 2.7 prompting guide walks through all three, with worked prompts, a mode-picker table, and the failure modes specific to each.

    Studio microphone in warm light

    Know which Wan you're prompting

    The three modes take different inputs, and each input changes what your words are for:

    Mode You provide Prompt's job Best for
    Text-to-video Words only Everything: subject, scene, camera, light Original scenes, b-roll, stylized shots
    Image-to-video A frame + words What happens next Product stills, portraits, meme frames
    Reference-to-video Identity refs + words Action and performance Recurring characters, persona content

    The rule that falls out of this table: the more the input carries, the less the prompt should describe. Text-to-video prompts run long; reference-to-video prompts run short. Writing a long, appearance-heavy prompt in reference mode doesn't add control — it adds contradictions.

    Text-to-video: build the scene in layers

    From a blank page, Wan 2.7 responds well to a layered structure — scene, then subject, then motion, then camera and light:

    A narrow Kyoto alley at dusk, paper lanterns glowing, light rain. A street cook fans skewers over charcoal, smoke drifting through the lantern light. Steam and smoke move with a slight breeze. Slow push-in from alley entrance toward the stall. Warm tungsten against blue dusk.

    What makes this prompt work is that every layer contains motion, even the environment — rain, smoke, breeze. Wan's text-to-video renders atmosphere well, but only when you put atmosphere in motion; a static description gets you a beautiful clip that feels like a wallpaper.

    Before: "A cozy cabin in the woods in winter, snow, warm lights, cinematic." After: "Snow falls steadily past a cabin window at night; woodsmoke curls from the chimney and wind pushes it sideways. A slow drift toward the glowing window. Warm interior light spilling onto blue snow." The fix: three moving elements (snowfall, smoke, camera drift) turn a postcard into a shot.

    Image-to-video: direct the next three seconds

    In i2v mode your frame already answers who, where, and what it looks like. Wan 2.7 wants continuation, and the prompts that land share a shape: first motion → secondary motion → camera behavior.

    From the image: she looks up from the book, closes it softly, and glances out the window. Curtain stirs in a breeze. Camera holds nearly static with a slight slow push.

    Three i2v rules that matter more on Wan than elsewhere:

    1. Bridge from the pose. The first motion must be reachable from the image — if the hands hold a book, the first beat involves the book, not a wave.
    2. Add one ambient motion (curtain, steam, traffic light change) to sell continuous reality.
    3. Restrain the camera. Wan i2v shots read best with holds and slow pushes; big prompted camera moves fight the fixed geometry of the source frame.

    Reference-to-video: identity from images, performance from you

    Reference mode is where Wan 2.7 earns its "References, Frames, and Voice" reputation. You load reference images to lock a character's identity, and — the distinctive part — Wan's reference pipeline supports voice-driven speech, so your recurring character can talk. Paired with Versely's voice cloning, that means a persona whose face and voice both stay consistent across an entire content series; the full workflow is covered in our Wan 2.7 reference-to-video voice clone guide.

    Prompting changes completely in this mode. Identity is handled; your words buy performance:

    He leans into frame like he's sharing a secret, taps the counter twice, and delivers the line with a slow grin. Small kitchen set, morning light, sitcom framing.

    Notice: zero physical description, all blocking and delivery. Direction verbs — leans, taps, delivers, smirks, sighs — are the vocabulary of this mode. For spoken content, keep each clip to one or two lines and write the delivery, not just the words: "delivers the line with a slow grin" produces visibly different footage than the bare sentence.

    Reference-mode failure to avoid: re-describing the character ("a man with a beard in a denim shirt…"). At best it's wasted words; at worst it drifts identity away from your references. If you want the character changed, change the reference set.

    One idea, three modes: a worked comparison

    Say the idea is "our mascot unveils the new flavor." Here's the same beat prompted per mode:

    • Text-to-video (no assets yet): "A round blue mascot with stubby arms stands at a tiny podium and whips a velvet cloth off a glowing soda can. Confetti drops. Static center framing, game-show lighting."
    • Image-to-video (you have the key art): "From the image: the mascot grabs the velvet cloth and whips it away, revealing the can. Confetti falls. Camera holds."
    • Reference-to-video (recurring mascot with a voice): "He drums the podium impatiently, whips the cloth away, gasps at the can, and says: 'We actually did it.' Game-show lighting, confetti."

    Same beat, three different prompts — and the reference version is the shortest while producing the most repeatable result. That's the Wan 2.7 pattern in miniature.

    Failure modes and fast fixes

    • Wallpaper syndrome (pretty but static) → no motion layers in a t2v prompt; add two moving environmental elements.
    • i2v warping in the first second → first motion unreachable from the pose; prompt a bridge beat.
    • Character drift in reference mode → appearance words in the prompt contradicting refs; strip them.
    • Flat line deliveries → you wrote dialogue without delivery direction; add a performance verb and a physical beat around the line.
    • Series doesn't feel like a series → set/light description varies per clip; freeze a one-sentence "set style" and reuse it verbatim.

    FAQ

    What are the three Wan 2.7 modes and when do I use each?

    Text-to-video builds scenes from words alone — best for original shots and b-roll. Image-to-video animates an existing frame — best for product stills and portraits. Reference-to-video locks a character's identity from reference images with voice-driven speech support — best for recurring personas and talking content.

    How is prompting Wan 2.7 reference-to-video different?

    Stop describing appearance entirely — references carry identity — and spend the prompt on blocking, action, and delivery. Direction verbs (leans, taps, grins) plus a fixed set-style sentence produce consistent, episodic content; appearance words only risk drifting the character.

    Can my Wan 2.7 character actually speak?

    Yes — Wan 2.7's reference pipeline supports voice-driven speech, and combined with Versely's voice cloning you get a persona with a consistent face and voice across a series. Keep clips to one or two lines and write the delivery, not just the dialogue.

    Why do my text-to-video results look like moving wallpapers?

    Static description. Wan renders atmosphere beautifully, but the prompt must put the atmosphere in motion — rain falling, smoke drifting, a slow camera push. Two or three explicit motion layers is the reliable fix.

    Is Wan 2.7 better than Kling or VEO for character content?

    For recurring characters with voice, Wan's reference mode is one of the strongest options; Kling leads camera-driven action and VEO leads native-audio cinematic shots. On a multi-model platform like Versely you can run the same character beat through each and compare, rather than committing on reputation.

    Pick the recurring character you've been meaning to build, load five reference images into Wan 2.7 reference-to-video in Versely, and prompt one short performance — a lean-in, two taps, one line. If episode one holds identity, you've got a series engine.