Guides

    Motion control needs a driving clip, not a paragraph

    A driving video is the input. Text cannot replace it. If you do not have a performance, you do not have a motion-control job.

    Versely Team6 min read

    Do not prompt "dance like this" without the clip. The product-versus-product pick against a preset-menu app is Versely vs Higgsfield.

    A driving video is the input. Text cannot replace it. If you do not have a performance, you do not have a motion-control job. You have an image-to-video job, or a text-to-video job, or a still that is not ready to move. Call it that. Do not open AI motion transfer and type choreography into a box.

    Motion control copies timing you already captured and puts it on a subject you supply. That is the only reason to use it. Pretty language about "dynamic movement" is how you spend a motion-control credit on an I2V result with extra steps.

    What the job is

    Two files:

    1. A driving clip — the walk, the dance, the gesture, the turn. Head, shoulders, and torso visible the whole way if the model is Kling-shaped. Subject fully in frame. You have the rights.
    2. A character still — framed like the clip. Full body if the motion uses a full body. A tight headshot against a full-body drive is the most common way this looks wrong.

    The model maps the pose sequence onto the still and keeps the clip's timing. Beats land where they landed. That is the product.

    DreamActor V2 does not even have a text field. Image URL, video URL, a toggle that trims the first second. Kling Video V3 Motion Control makes prompt optional, 0–2500 characters, for wardrobe and backdrop and mood — not for the movement. The movement is video_urls. If you write a camera essay into that box, you are arguing with a file.

    Video editing is where the result goes after: trim, composite, caption. It is not where you invent the dance.

    The decision

    You have You do not have Job
    A performance clip + a character still, framing matched A need to invent new choreography Motion control
    A still you love Any clip of the move Image-to-video. Describe motion. Accept invented timing
    A paragraph about a dance A still and a clip Not motion control. Not yet. Film or find the clip, or drop the job
    A clip you do not have rights to Not a job you should run
    A talking-head line that must match a voice A body performance Lipsync / native audio. Wrong family
    A product that must rotate 180° A turntable clip First-last frame, not motion control

    If the exact timing matters — music, a gesture that has to hit a word, a mascot that has to reuse last week's wave — motion control is the only family that takes timing as an input. I2V will invent a wave. It will not invent that wave.

    If the exact timing does not matter, you are overbuying. I2V a locked still and describe "she turns to camera, slow." Cheaper, fewer files, no framing fight.

    You cannot prompt the clip into existence

    The sentence people try: "A woman dances like the reference, energetic, matching the beat, cinematic handheld."

    There is no reference. There is a hope. The model will generate a dance. It will not generate the dance you saw on TikTok, the dance your founder did in the hallway, or the dance in the brief's unlinked Drive folder. Motion-control endpoints that require video_urls will simply fail. Endpoints that make the clip optional will silently become I2V.

    Film it. Phone, plain wall, full body, one take, keep them in frame. Or use a clip you own. Then the still. Then generate. The motion transfer path is those two files and a generate. Guidance text, if the model offers it, is "now wearing the red jacket, studio grey," not "then she does a pirouette."

    Kling's two switches are the other contract: character_orientation (whose facing wins, image or video) and background_source (whose backdrop wins). Set them. Do not try to set them in prose.

    Framing is the other way to not have a job

    A driving clip of a full-body dance plus a still of a face is not a motion-control job. It is a mismatch. Kling's schema wants head, shoulders, and torso on both uploads, aspect between 2:5 and 5:2. Crop before you upload. Do not ask the prompt to "zoom out."

    DreamActor's default trims the first second. If the opening beat is the beat, turn that off. The clip is the brief. Settings that throw the brief away are not a style.

    What this is not

    Not I2V. I2V invents motion from a still and a description. Use it when invented motion is fine.

    Not R2V. R2V recasts identity in a new scene. The scene's motion is still invented unless you also have a video reference on a model that takes one — and even then, "video reference" is not always "pose transfer." Read the endpoint. Motion-control pages are labelled motion-control.

    Not a hook pack. A hook is a first-three-seconds promise. A driving clip is a performance. Do not motion-control ten openings unless you have ten performances.

    Not "make it dance." That is a prompt. Film the dance.

    FAQ

    Can I describe the dance in text and skip the driving clip if the model is good enough now?

    No. Better models still need the input the family is built on. Without a clip you are in a different family. Call I2V. Do not call it motion control because the output happens to move.

    Does the driving clip have to be my body?

    It has to be a performance you have the rights to, framed like the still. Your body is the usual source because it is clean and you own it. A stock clip you licensed can work. A viral clip you did not license is not a workflow.

    Why did my transfer keep the driving clip's background?

    Because the default on Kling is that the video wins the backdrop unless you set background_source to the input image. Treat subject-level jobs as subjects: composite later in video editing if you need a new room. Do not fight the backdrop with a paragraph.

    When should I use motion control for a product instead of a person?

    When the product has a performance — a hand demo, a pour, a wearable in motion — and you have filmed that performance. A bottle that must simply turn is first-last-frame. A bottle that must be handled the way your founder handles it is a driving clip of the founder plus a still of the bottle in the same framing, or a person-plus-product transfer. No clip, no transfer.