Workflows

    Lock the still before you spend on motion

    Image-to-video and R2V are how you stop paying for identity retries. Text-to-video is how you start them.

    Versely Team7 min read

    If the still is wrong, every second of video is waste. That is not a taste argument. It is a credit argument. Video generation spends the expensive meter on motion, physics, and (on some models) audio. It does not spend that meter on discovering whether the face, the bottle, or the room is the one you meant. When you start from a paragraph, you find out after the render. When you start from a locked still, you find out at image prices.

    A character sheet before the first video credit is the sheet itself — face, wardrobe, palette, approved in stills. This post is the routing rule that sits on top of it: image-to-video and reference-to-video are how you stop paying for identity retries. Text-to-video is how you start them.

    Video credits are a terrible identity debugger

    A text-to-video job has to invent the subject and move it. Those are two jobs. When the output comes back with the wrong jawline, the wrong label, or a kitchen that is not your kitchen, you did not fail at motion. You failed at identity, and you paid the motion rate to learn it.

    Retrying the same prompt does not lock the subject. It samples another guess. Same prompt, new noise, new face. That is why a "character" that only lives in a paragraph drifts across a sequence, and why a product that only lives in adjectives will not match the pack shot. The model is not being stubborn. You never gave it a contract.

    The cheap place to argue about identity is a still. Iterate there until someone can point at a frame and say "that one." Then spend video credits on the only remaining job: making that picture move.

    Two locks, two contracts

    Once the still exists, pick the endpoint that matches what you are trying to keep.

    Image-to-video treats the still as the first frame. Composition, background, lighting, and product placement are already decided. The model invents motion forward from that picture. Use it when the frame is the deliverable — a pack shot that needs a slow push, a character in a room you already designed, a hook poster that has to load as that exact image. AI image-to-video is that path.

    Reference-to-video treats the still as a casting call. Identity travels; the scene is new. Use it when the same person or SKU has to appear in rooms you do not have a photo of. Do not feed a finished hero frame to R2V and ask it to "just add motion." It will restage. The split, in one sentence each, is image-to-video vs reference.

    Do not mix both on one shot. Do not "fix" a bad I2V take by throwing the last frame at R2V. Pick the contract, then honour it.

    Reusable assets belong in the sheet, not in the prompt. Set up reusable characters and products once, then call them by key. That is how a face survives ten scenes without being re-described ten times.

    What "locked" actually means

    A still is not locked because you generated one. It is locked when the things that would force a regen are already decided:

    • Identity. Face, body, wardrobe. If the character sheet is not signed, do not animate.
    • Product. Label, colourway, angle. If the bottle does not match the pack, every second of motion is a recall.
    • Room. Palette, practical lights, the wall the camera will see. A new room is a new still, not a new paragraph on the same video job.
    • First-frame job. On short-form, the opening still is the poster. If that poster is wrong, motion cannot save the hold.

    The images-first board is the same doctrine at sequence scale: storyboard brand videos with AI images first, approve cheap, animate only what is locked. This post is the per-shot version of that board. One still. One yes. Then motion.

    When text-to-video is still the right first move

    T2V is not banned. It is banned as an identity tool.

    Use it when there is no subject to lock: weather, particles, abstract motion, a cutaway where "a kitchen at dusk" is the whole brief and nobody will match it to a SKU. Use it for exploration, when you do not yet know what the scene should look like and surprise is the point. Then stop. Take the frame you would actually ship, still it, and switch to I2V or R2V for the takes that have to match each other.

    The failure mode is using T2V for the hero because it felt faster, then spending the rest of the day regenerating because the product drifted. Faster to first file is not faster to locked file.

    The credit test

    Before you hit generate on a video job, answer one question: if this still were printed on a box, would we ship the box?

    If no, do not buy motion. Fix the still. If yes, write a motion prompt that describes only movement — camera, subject action, atmosphere — and leave identity out of the paragraph. Identity is already in the frame. Repeating it in text is how you invite the model to renegotiate.

    A regen of motion on a locked still is a new take of the same shot. A regen of a T2V paragraph is a new casting session. Those are not the same cost, even when the credit line looks similar, because one of them has a chance of matching what you already approved.

    What you actually send

    The input that wastes credits is a paragraph plus "make it consistent." The input that spends credits once is:

    • a still (or a small reference set) someone has signed
    • a motion line that names camera and action only
    • the endpoint that matches the contract — I2V if this frame is the start, R2V if this face has to appear in a new room

    If you cannot point at the still in the job, you are still in exploration. Exploration is allowed. It is not a hero take, and it should not be billed or scheduled as one.

    FAQ

    Can I skip the still if the model is good at characters now?

    No. Better models reduce drift. They do not create a contract. If the character has to match itself tomorrow, you still need a still (or a reference set) that tomorrow's job can see. Prompt memory is not that still.

    Is a last-frame hold enough to lock identity?

    Only for the clip that ends on it. The next shot does not inherit that hold unless you feed the frame back in as I2V or as a reference. Saving the render is the durable copy. Hoping the next text prompt remembers the face is not.

    What if the still looks right and the video still drifts?

    Then you have a motion problem, not an identity problem — or you picked R2V when you needed I2V. Check the contract first. If I2V is drifting off a locked frame, shorten the shot, simplify the action, or change model. Do not "fix" it by rewriting the character into the prompt.

    Should product shots always be I2V?

    When the composition is the ad, yes. When the product has to appear in a new location, R2V from a product still is the lock. Text-to-video is how you get a bottle that is almost yours.