Guides

    Directing gaze and micro-expression on a face

    Emotion words return a stereotype. Eyeline, blink timing and independent brow movement are directable as physical actions — and here is where prompting stops.

    Versely Team9 min read

    A prompt that says "she looks sad" returns the average of every sad face the model was trained on: brows raised in the middle, mouth corners down, eyes slightly wet, held perfectly still for the whole clip. It is legible and it is dead. Nobody believes it, because real faces do not hold one expression. They arrive at one, sit in it briefly, and leak something else on the way.

    The fix is not a better adjective. It is to stop writing emotions and start writing the physical events that produce them: where the eyes point, when they move, when the lids close, what the brow does a beat before the mouth. All of that is writable. The emotion is the viewer's conclusion, not the model's instruction.

    Why emotion words underspecify

    An emotion word names a category, and the model renders something near the middle of that category's distribution. That is why "angry" returns roughly the same furrowed, faintly theatrical face across models and seeds. The token is a label, and labels resolve to stereotypes.

    It is the same mechanism that makes "cinematic lighting" return a look rather than a light source. Naming one physical thing — where the lamp sits, where the eyes point — narrows the output further than three adjectives that all gesture at the same fuzzy region.

    Physical descriptions escape this because they name an executable state. "Eyes lowered to the mug in her hands" is a geometry instruction. "Brow held level while the mouth tightens" specifies two regions independently, which is exactly what a stereotype cannot do — a stereotyped face moves as one unit.

    Working rule: one emotion word is allowed as context at the front. Everything after it should be physical.

    The four channels you can actually write

    Eyeline. Two decisions: where the eyes are pointed, and when that changes. "Looking at the camera" and "looking just past the camera, screen left" produce visibly different shots; the second is what interview footage looks like. Naming a target inside the frame beats naming a direction, because the model can resolve an object but has to guess at "down". "Eyes fixed on the mug in her hands" lands more reliably than "looking down".

    Lids and blinks. A held stare with no blink turns uncanny in about three seconds, and models default either to no blink at all or to a metronome. Direct it explicitly: "a single slow blink as she finishes the sentence." Half-lidded, squinting against a bright source, and eyes widening are all lid instructions, not emotions.

    Brow and mouth, separately. The tell of a real performance is that the two disagree for a moment. "The brow stays level while the mouth pulls into a small closed smile" is a physical description of what a viewer will call wry. Write them as two clauses so the model has to resolve them independently.

    Head and neck. Small amounts. A few degrees of tilt, a chin drop, a turn that leads or lags the eyes. Eyes first then head reads as noticing something. Head first then eyes reads as having been told to look.

    You wrote Write instead
    sad eyes lowered to the table, one slow blink, mouth relaxed and slightly open, chin dropped a few degrees
    suspicious eyes stay on the other person while the head turns a few degrees away, brow level, jaw set
    relieved shoulders drop on a long exhale, eyes close for a beat, then open toward the window
    trying not to laugh lips pressed together, cheeks lifting, creases at the outer corners of the eyes, one glance off-camera
    deadpan eyes stay on the lens, no blink through the line, no brow movement, mouth neutral

    That last row is the one to test before you trust it. Deadpan is the register where doing nothing is the direction, and models push against it the same way they push toward adding camera movement when you asked for none. The deadpan prompt set has generated clips attached, which is a faster way to see how the phrasing behaves than writing your own from scratch.

    Timing is the part text struggles to carry

    Speed and easing get re-estimated on every generation. "A slow blink" is interpreted fresh each run, which is why a performance that works on take three is not reproducible from the same words on take four. Two things help.

    Shorten the clip. A three to five second generation has fewer events to schedule and holds a directed performance far better than an eight second one, where the model starts inventing beats to fill the time. Short clips are also where temporal consistency is strongest in general, so you are buying two things at once.

    Then place the beats explicitly where the model supports it. Veo 3.1 accepts per-window prompting, and Google's own Veo 3.1 prompting guide documents the format along with a prompt skeleton that leads with cinematography rather than subject. Both details are worth copying: shot size first, performance second, and each beat in its own window.

    Medium close-up, 85mm, the camera is entirely motionless.
    [00:00-00:02] She is looking down at the mug in her hands. Brow level, mouth neutral.
    [00:02-00:04] Her eyes lift to the lens. One slow blink on the way up. The mouth
    stays closed; only the outer corners of the eyes move.
    Warm practical lamp at frame left, ambient room fill.
    

    "The camera is entirely motionless" is doing real work in that prompt. Omitting movement language does not give you a locked-off frame — the prior runs toward adding a drift or a slow push, and you have to say no to it in words.

    Checking the take

    Three things, in this order, before you accept a face:

    1. Eyeline continuity. Do the eyes stay on their target between the moments you directed a change, or do they wander?
    2. Blink count. Zero blinks in five seconds is a tell. So is a blink every second.
    3. Whether the brow and mouth moved together. If they did, you got the stereotype back, and the prompt needs the two clauses split further apart.

    The fourth failure only shows up on longer clips: reversion. The face drifts back toward a neutral stereotype in the last second or two, because the directed state was a starting condition rather than a sustained one. Cutting before the reversion is usually cheaper than re-rolling.

    Where prompting stops and transfer starts

    Prompting a face works for a beat. One eyeline change, one blink, one expression arriving. It stops working when you need a sustained performance: a whole line of dialogue with human timing, hands included, holding through a shot.

    At that point the right move is transfer rather than description. You perform it, and the model maps the performance onto your character.

    • Runway's Act-Two takes a driving performance video plus a character reference and tracks head, face, body and hands onto that character. Runway's documentation notes that when the character is a still image, the gesture-control setting has to be enabled to get body and hand motion instead of head movement alone.
    • In the Versely catalog, Dreamactor V2 is the face-reenactment route: facial expressions and head movement transferred from a driving video onto a reference image. Motion transfer is the tool surface, and transferring a motion onto a photo is the agent route if you would rather describe the job than choose a model.
    • If the constraint is dialogue rather than expression, the lipsync path exposes emotion, expression and talking style as parameters rather than prose, which is more repeatable than adjectives across a batch. Three ways to make a face talk covers which route fits which job, and the ranked lipsync view shows credit costs beside each option.

    The threshold is simple enough to apply without thinking about it: if you can name the performance as one or two physical events, write it. If describing it takes three sentences of timing, film yourself doing it and transfer it instead.

    FAQ

    Do emotion words ever help?

    As context, yes. Leading with one word gives the model a region of expression space to work inside, and the physical clauses after it then narrow that region. What fails is an emotion word carrying the whole instruction, because there is nothing left to disambiguate which of a thousand versions of "worried" you meant.

    Why does the face go blank halfway through the clip?

    Because the directed expression was treated as a starting state, not a sustained one, and the model relaxes toward its neutral prior as the clip runs. Shorter generations help most. Failing that, direct a second event later in the clip so there is something to arrive at, or cut before the drift becomes visible.

    Can I direct gaze the same way in a still image?

    Mostly. Eyeline, lid position, brow and head angle all transfer directly to image prompts, since they describe a single frame. Blink timing and any "then" clause do not — a still has no timeline, so anything phrased as a sequence collapses into one ambiguous pose.

    Is performance transfer always better than prompting?

    No. Transfer costs you a shoot, even if the shoot is a webcam take, and it locks the performance to whatever you recorded. Prompting stays faster for single beats, reaction shots, and anything you want twenty variants of. Transfer wins specifically on sustained dialogue and on hands, which prompting handles worst.