Two things go in: a clip that carries the movement, and an image of whoever should be doing it. The model reads pose and timing from the first and re-renders it wearing the second. Choreography, gesture rhythm and the beat of a gesture survive; wardrobe, lighting and set come from the character image and the prompt.
It is the only mode where you can specify motion exactly rather than describe it. "She spins twice and points at the label" is a paragraph of prompt with an uncertain outcome; a two-second clip of someone doing it is unambiguous.
The practical constraint is that the driving clip has to be readable. Full-body framing with clean separation from the background transfers well; crowded frames, extreme close-ups and heavy occlusion do not, because the pose the model extracts is only as good as what it could see.
In practice
- The driving clip sets the length and the beat — you are casting a performance, not editing one.
- Body proportions between driver and subject affect fidelity; wildly different builds produce odd limb work.
- Camera movement in the driving clip may or may not transfer, depending on the model.
Motion-control models
Catalog entries that transfer movement from one clip onto another subject. 7 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Kling Video V3 Pro Motion Control | Kling | Video |
| Kling Video V3 Standard Motion Control | Kling | Video |
| DreamActor V2 | ByteDance | Video |
The mistake to avoid
Expecting facial performance to come across. Most motion transfer is body-level; expression and lip movement usually need a separate lipsync pass.
Where you will run into it
- AI Video Generator — Text-to-video, image-to-video, and story chaining in one studio.
- AI UGC Video Generator — UGC ads at the speed and price of a prompt.
Related terms
Reference-to-video
Reference-to-video means the model builds a clip from separate images of a subject it can place anywhere. Those images are an identity, not the first frame.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Character consistency
Character consistency meaning: whether the same person, mascot, or product still looks like itself across separate generations.
Camera control
Camera control is the set of inputs that decide how the virtual camera behaves during a generated clip — whether it pushes in, orbits, pans, or stays locked off.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.