Generation modes

    Motion control

    Also called Motion transfer.

    Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.

    Two things go in: a clip that carries the movement, and an image of whoever should be doing it. The model reads pose and timing from the first and re-renders it wearing the second. Choreography, gesture rhythm and the beat of a gesture survive; wardrobe, lighting and set come from the character image and the prompt.

    It is the only mode where you can specify motion exactly rather than describe it. "She spins twice and points at the label" is a paragraph of prompt with an uncertain outcome; a two-second clip of someone doing it is unambiguous.

    The practical constraint is that the driving clip has to be readable. Full-body framing with clean separation from the background transfers well; crowded frames, extreme close-ups and heavy occlusion do not, because the pose the model extracts is only as good as what it could see.

    In practice

    • The driving clip sets the length and the beat — you are casting a performance, not editing one.
    • Body proportions between driver and subject affect fidelity; wildly different builds produce odd limb work.
    • Camera movement in the driving clip may or may not transfer, depending on the model.

    Motion-control models

    Catalog entries that transfer movement from one clip onto another subject. 7 of the 296 models in the Versely catalog qualify.

    The mistake to avoid

    Expecting facial performance to come across. Most motion transfer is body-level; expression and lip movement usually need a separate lipsync pass.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.