Two things go in: a clip that carries the movement, and an image of whoever should be doing it. The model reads pose and timing from the first and re-renders it wearing the second. Choreography, gesture rhythm and the beat of a gesture survive; wardrobe, lighting and set come from the character image and the prompt.
It is the only mode where you can specify motion exactly rather than describe it. "She spins twice and points at the label" is a paragraph of prompt with an uncertain outcome; a two-second clip of someone doing it is unambiguous.
The practical constraint is that the driving clip has to be readable. Full-body framing with clean separation from the background transfers well; crowded frames, extreme close-ups and heavy occlusion do not, because the pose the model extracts is only as good as what it could see.
In practice
- The driving clip sets the length and the beat — you are casting a performance, not editing one.
- Body proportions between driver and subject affect fidelity; wildly different builds produce odd limb work.
- Camera movement in the driving clip may or may not transfer, depending on the model.
Motion-control models
Catalog entries that transfer movement from one clip onto another subject. 7 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Kling Video V3 Pro Motion Control | Kling | Video |
| Kling Video V3 Standard Motion Control | Kling | Video |
| DreamActor V2 | ByteDance | Video |
| Wan 2.2 14B Move | Wan | Video |
The mistake to avoid
Expecting facial performance to come across. Most motion transfer is body-level; expression and lip movement usually need a separate lipsync pass.
Where you will run into it
- AI Video Generator — Text-to-video, image-to-video, and story-to-video in one place.
Related terms
Reference-to-video
Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Character consistency
Character consistency is whether the same person, mascot or product still looks like itself across separate generations.
Camera control
Camera control is the set of inputs that decide how the virtual camera behaves during a generated clip — whether it pushes in, orbits, pans, or stays locked off.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.