Guides

    Your image-to-video clip barely moves

    A clip that reads as a photo with drift is an underspecified motion brief. Write one subject action and one camera move, then fix the source frame.

    Versely Team9 min read

    The model did not "fail to animate." It did exactly what an underspecified image-to-video brief asks: hold the still, then invent a little life so the file is not a JPEG. What you get back is a photograph with a slow push-in nobody requested, or a subject whose hair twitches while the body stays posed. That is the default, not a weak sampler.

    Image-to-video already knows what the frame looks like. The prompt's only job is to say what changes. If you re-describe the room, the wardrobe and the lighting, you have written a stills prompt on a motion model. The motion budget gets filled with whatever the training prior likes, which is a gentle drift. The fix is a brief with one subject action and one camera move, plus a source frame that can actually survive that action.

    The prompt is a motion brief, not a scene description

    The still has already locked identity, wardrobe, set and grade. Repeating them in the prompt competes with the pixels. On every image-to-video model that is actually animating a frame, the prompt's job is the motion and action, not another pass at the scene.

    Write two clauses. Only two.

    One subject action. A single physical verb the body can complete in the duration you picked. "She turns her head toward camera and smiles." "He picks up the mug and takes a sip." "Steam rises from the cup; her hands stay still." Not "she stands up, walks to the window, opens it and looks out" on a five-second clip.

    One camera move, in film vocabulary, at the end. "Slow push-in toward her face." "Locked-off, tripod, the camera does not move." "Gentle pan left with her eyeline." One move per shot is the reliable brief. Two named moves ("push in then pan right") make the model attempt both and land neither. Put the camera last so it is direction on top of an already-stated action, not a second scene description.

    A shape that holds up:

    [WHO / WHAT is already in the still] [ONE ACTION, present tense].
    Secondary motion: [one supporting detail, or none].
    Camera: [ONE named move, with speed].
    

    Examples, against the same source still of a woman at a kitchen counter:

    Brief What you usually get
    "Cinematic, beautiful morning light, professional, 4K, she looks natural" A living still. Maybe a drift.
    "She is happy and the camera is cinematic" Smile plus a wander.
    "She turns to the kettle, then pours, then sips, then smiles while the camera pushes in and orbits" A compromise of all of it, often with less net motion than one clean action
    "She turns her head toward camera and smiles. Steam rises from the kettle. Camera: locked-off on a tripod." A readable action. The still stays the still.

    Silence on the camera is not a locked-off shot. Unspecified camera behaviour gets filled with default drift or a slow push-in. If you want stillness, say it: "static, locked-off, tripod." The full phrasing for that case lives in how to get a genuinely locked-off shot. If you want movement, name it the way an operator would. A working glossary of the moves models actually parse is in camera movement prompts.

    Camera control and subject language compete. A fast orbit around a person who is also running is two hard jobs. Give the motion to one of them. A locked camera with a subject who walks through the frame has somewhere to put the motion budget. A locked camera with a subject who also stands still is a photograph, and photographs get drift.

    Turning the motion dial up is the wrong rescue

    Some models expose a motion level as low, medium or high. It is an amount dial, not a verb selector. High means more change per frame across subject, background and camera. It will not make "she sips the coffee" happen if you never wrote that she sips the coffee. It will add drift, background crawl and a face that works harder than the action.

    The glossary pitfall is the whole failure mode of this post: turning motion up because the clip felt static, when static meant the prompt named no action. The dial then fills the gap with the default wander.

    Where the dial exists, use it after the prose is specific:

    • Low once the action is named, for talking heads, product and anything with type in frame.
    • Medium for a single clear gesture plus a slow camera.
    • High only for atmosphere and abstract b-roll, where anatomy is not under inspection.

    Where the dial does not exist, the verbs are the dial. "Drifts", "gently", "almost still" produce a living photograph. "Turns", "reaches", "steps", "pours" produce a shot. Stacking four of those verbs produces a mess, not more energy.

    Duration has to fit the action. A head turn on a ten-second clip pads with drift after the turn is over. A walk across the room on a three-second clip becomes a teleport. Size the action to the length, or cut the length to the action. Do not ask the model to "just make it more alive" for the unused seconds.

    Fix the source frame before you rewrite the prompt again

    A motion brief cannot animate information the still does not contain. If the subject is a soft smear in the corner of a wide interior, the model has nothing to articulate. It will hold the smear and breathe a little, which is exactly the "barely moves" clip.

    Before another generation, look at the still the way the model has to: no memory of the object, just pixels.

    • Crop to the subject and view at 100 percent. If that crop is soft, noisy or blocky, that is the resolution the sampler is working with. Upscale the still first, or reshoot it. The input rules for this are in garbage reference in.
    • The first frame of the finished clip is this photo. Do not feed an extreme angle you assumed the model would "correct" as soon as the camera moved. If you asked for a push-in, the still should already be a frame you are willing to open on.
    • A clean cutout is not always a better plate. Stripping the background also strips contact shadows and scale cues. A subject that has to walk, sit, or put a mug down needs those cues, or the body stays frozen because the model cannot ground the next pose.
    • Aspect ratio comes from the still on a large part of the catalog (Vidu, Pixverse, Wan image-to-video among them). Writing "16:9 widescreen" in the prompt does nothing if the upload is square. Crop the source to the delivery ratio before you animate.

    If the still is a posed hero (square on, hands folded, looking down the lens), "natural candid motion" is a fight with the pose. Either pick a still that already implies the action (weight on the front foot, hand near the mug, head already mid-turn) or write an action that a posed body can actually do (a smile, a blink, a small head tilt). Asking a mannequin still to sprint is how you get a mannequin with drift.

    A test you can run in one session

    Lock the still. Change only the brief.

    1. Zero-motion control. "The subject holds still. Camera: locked-off on a tripod." If even this drifts, the model is filling silence and you need the explicit stillness language, not more adjectives.
    2. Subject only. One action, locked camera. If this still barely moves, the action is too small, too slow, or impossible from the pose.
    3. Camera only. "The subject holds still. Camera: slow push-in toward the face." If this is the only take that feels alive, you have been asking for camera energy and not writing it.
    4. The real brief. One action plus one camera move. That is the production prompt.

    Three takes of (4) is a normal reroll budget. If none of them move, stop swapping models and fix the still or the verb. The best image-to-video models page is the live ranking; a dead still plus a stills prompt will look static on every row of it. If a take finally moves, inspect what moved. Camera-only drift means the action is still missing. A warped body means the verb was too big for the still. Cut the verb down. Do not add a second one.

    FAQ

    Why does "make it cinematic" never produce a real camera move?

    Because it names a look, not a behaviour. Models already have a prior for "cinematic" (grade, shallow depth, a slow push). You are requesting the average of every dramatic clip they have seen, which is drift with nicer colour. Name the move: dolly, pan, orbit, locked-off.

    Should I put the camera instruction first or last?

    For a motion brief, put the action first and the camera last, as one named move. The still has already set the frame; the last clause is how that frame is allowed to change. The exception is a true locked-off shot, where leading with stillness stops the model committing to a wander before it reads the rest. Do not do both: a paragraph that opens "locked-off" and ends "slow push-in" is two camera briefs.

    Does a higher motion level replace the action clause?

    No. Motion level scales how much change occurs. It does not choose which change. An empty brief plus high motion is how you get boiling backgrounds and a face that shimmers while the hands never reach for the prop. Write the reach, then use the dial to keep it contained.

    The subject moves in the still's pose but will not walk. Is that the model?

    It is the still. A seated, cropped, or tightly posed figure does not contain the geometry of a walk. Either generate a wider still that includes the legs and a path, or write an action a seated body can do. Image-to-video animates the photograph you fed it. It does not cast a new shot.