Kling Motion Control Prompting Guide
Teach Kling Motion Control as image plus driving clip. Prompt scene and style, not choreography. Matches Video vs Image, failure modes, Versely path.
Kling Motion Control is not another text-to-video prompt dialect. It is a transfer job: a character still plus a driving clip. The video owns timing, limbs, and often camera. Your text owns environment, wardrobe notes you must keep, and style. On Versely that job lives next to plain Kling T2V/I2V, and the named agent path is transfer a dance or motion onto my photo. The tool-shaped door is AI Motion Transfer.
This guide owns motion-control teaching. It is deliberately distinct from the Kling 3.0 prompting guide, which covers shot briefs, multi-shot, and Turbo. If you only have adjectives and no driver, stay on T2V/I2V. If timing is the content, bring a clip.
The contract: still + driver, not a dance paragraph
Motion control maps a performance onto a subject. Kling (and the Versely rows that wrap it) expect:
- Character image - who appears. Face readable, body proportions clear, framing that matches how much of the body the move needs.
- Driving / reference video - what happens. Continuous shot, moderate speed, head and body visible, minimal occlusion.
- Optional prompt - scene, lighting, style, and (in one orientation mode) camera treatment.
- Character orientation - Matches Video or Matches Image (see below).
Official Kling and provider docs agree on the scarce asset: the driving clip. You typically need at least about 3 seconds of usable continuous motion. Matches Video often allows longer drivers (about 30s); Matches Image is shorter (about 10s). Trust the live row for caps.
The driver is also the rights problem. You are copying a performance. Film it, license it, or get consent. Do not treat a stranger's dance as free puppeteering.
Prompt environment and style, not choreography
The model already has the dance, gesture, or delivery in the video. Rewriting eight bars of choreography in prose fights the driver and wastes credits.
Prompt for:
- Place and time of day
- Light direction and mood
- Wardrobe or prop notes that must survive the transfer
- Style (photoreal, illustration look, brand grade)
- In Matches Image: camera moves Kling documents for that mode (zoom in/out, camera up/down, fixed)
Do not prompt for:
- Step-by-step dance instructions the clip already contains
- "Exact TikTok trend" without attaching the trend's move as video
- A second character's motion when only one subject is transferrable
Worked prompt 1 (Matches Video: performance accuracy)
SCENE: Clean studio with soft gray seamless, cool key from camera left.
SUBJECT HOLD: Keep the red bomber jacket and white sneakers from the still.
STYLE: Photoreal social cut, natural skin, no stylization.
CONSTRAINTS: No on-screen text, no extra people, no prop invent.
Attach a full-body still of your character and a continuous full-body driving clip of the routine. Orientation: Matches Video. Do not narrate the footwork.
Worked prompt 2 (Matches Image: keep photo facing, prompt camera)
SCENE: Same brick loft bedroom as the still, late afternoon window light.
CAMERA: Slow zoom in, then hold. Fixed height, no orbit.
SUBJECT HOLD: Keep her short black bob and green wool sweater from the image.
STYLE: Soft natural grade, shallow depth.
CONSTRAINTS: No new furniture, no subtitles, no wardrobe change.
Use when the still's facing and framing must stay boss, and the clip still supplies body and expression timing. Cap the driver to what Matches Image allows on your row.
Worked prompt 3 (product-demo gesture, not a dance)
SCENE: White product table, soft overhead softbox, subtle reflection on matte surface.
ACTION CONTEXT: Hands demonstrate the bottle pour only as in the driving clip.
SUBJECT HOLD: Same frosted citrus bottle label from the still.
STYLE: Clean e-commerce hero, photoreal.
CONSTRAINTS: No logo invent, no second bottle, no text overlays.
Here the scarce asset is a clean hand-performance clip you filmed. The prompt only protects the set and the SKU.
Matches Video vs Matches Image
| Mode | Who owns facing / orientation | Best for | Typical driver length |
|---|---|---|---|
| Matches Video (default on many UIs) | Driving clip: pose, facing, often camera follow | Complex full-body choreography, accurate performance | Longer (often up to ~30s) |
| Matches Image | Character still keeps orientation; motion and expression still follow the clip | Protecting photo composition; prompting camera (zoom, tilt cues) | Shorter (often up to ~10s) |
Pick deliberately. If performance accuracy is the goal, Matches Video usually holds up better on hard moves. If the hero is a carefully framed portrait and you only need gesture timing plus a prompted push-in, try Matches Image. On Kling VIDEO 3.0, facial element binding for stronger identity is documented for video orientation paths; do not assume it on every Matches Image run.
Framing match is the first quality gate
Official Kling notes repeat the same rule creators learn the hard way: match full-body to full-body and half-body to half-body. A close-up face still plus a wide dance driver is the classic failure. Leave room in the still for limbs to travel. Prefer one readable subject in the driver; multi-person clips often transfer whoever occupies the largest area.
Good driver habits: single continuous shot, moderate speed, head and body unobstructed, stable light with background separation, and clear humanoid proportions when limbs matter.
Failure modes
| Failure | Likely cause | Fix |
|---|---|---|
| Limbs melt or pose snaps | Framing mismatch (close-up still + full-body driver) | Rematch body crop; re-shoot still or cut driver to matching scale |
| Identity drift on turns | Weak face lock / no element bind where supported | Cleaner face refs; bind facial element on Matches Video when the UI offers it |
| Occlusion glitches | Hands, props, or hair covering face/torso in the driver | Re-film with clear silhouette; avoid crossed arms over face |
| Take cuts mid-move | Cuts, whip pans, or shot changes in the driver | One continuous take; trim to valid continuous segment (≥~3s) |
| Shorter output than upload | Complex or very fast action; model keeps only valid motion | Slow the performance; simplify; accept shorter billable extract |
| Wrong person puppeteered | Two+ people in the driver | Solo driver, or accept largest-on-screen selection |
| Prompted dance ignored | You described choreography instead of attaching it | Attach the clip; shrink prompt to scene/style |
| Camera fights the body | Matches Video with a prompted orbit that conflicts | Use Matches Image for prompted camera, or let Matches Video own camera |
| Rights hangover | Scraped influencer dance without consent | Own the performance, license it, or get written permission |
Versely rows (catalog snapshot)
Prefer AI Motion Transfer and the agent job over a T2V workaround. Published pages:
- Kling Video V3 Standard Motion Control - 13 credits/s, up to 720p
- Kling Video V3 Pro Motion Control - 17 credits/s, up to 1080p
Both need a starting image and a video. Length follows the driver. Confirm the live composer total. Sibling: motion transfer is a named agent job.
Skip the dialect: brief the Versely agent
You can learn Matches Video vs Matches Image. On Versely you can also brief the agent instead of writing prompts: attach the still and the driver, name the job ("transfer this motion onto my photo"), state orientation preference if you care, list scene/style constraints, platforms, and a spend ceiling.
Example job brief:
Job: motion transfer (not plain I2V).
Attachments: character still (full body) + driving clip (solo, continuous, 6s).
Orientation: Matches Video.
Prompt constraints: gray seamless studio, keep red jacket, no text, no extra people.
Output: 9:16 social cut. Budget: stop if quote exceeds my cap; show credits before run.
A hand-tuned transfer still wins for one hero take. A job brief wins when you want the agent to route to transfer a dance or motion onto my photo, pick Standard vs Pro, and keep you from burning a motion-control meter on an adjectives-only generate.
After the take: edit, post, collections
When the transfer is close enough:
- Trim to the beat you need. Do not ask the model to be your NLE.
- Composite onto a cleaner background if the pass is subject-strong but set-weak (common on transfer rows).
- Burn mute-proof captions with /tools/ai-caption-generator or /free-tools/burn-captions when the plate is silent or the audio from the driver is not what you will ship.
- Upload or schedule where Versely already connects your accounts; otherwise export the mp4 and post natively.
- Save keepers into a Versely collection: the winning still, the rights-cleared driver, and the scene prompt that survived. Reuse the same driver across characters when that is the campaign plan.
Honest limit: one driver plus one still is one performance copy. Board separate generates for separate outfits or locations, then stitch if you need a sequence.
FAQ
Is Kling Motion Control the same as Kling 3.0 text-to-video?
No. T2V invents motion from language. Motion Control copies motion from a driving video onto a character image. Use the Kling 3.0 prompting guide for shot briefs; use this page for transfer technique.
What should I put in the prompt?
Environment, light, style, wardrobe hold, and constraints. In Matches Image, add camera treatment. Do not rewrite the choreography the clip already owns.
Matches Video or Matches Image?
Matches Video for complex performance accuracy and longer drivers. Matches Image when the still's orientation must stay boss and you want prompted camera. Test both on a short cut of the same driver before you burn a long take.
Why did my output come out shorter than the upload?
Kling extracts valid continuous motion. Cuts, extreme speed, or messy framing shorten the usable segment. Keep at least about three seconds of clean action; trim and simplify before you re-roll.
Can I use any dance video from social media?
Only if you have rights or consent. The driving clip is a performance asset. Scraping a stranger's routine into a commercial or brand post is a legal and trust problem, not a prompting trick.
Where do I run this on Versely?
Open AI Motion Transfer or the AI video generator with a motion-control row selected, or ask the agent to transfer a dance or motion onto my photo with both files attached. Check Standard (13 cr/s) vs Pro (17 cr/s) on the model pages before you confirm.
Takeaway
Treat Kling Motion Control as image + driving clip. Prompt the set and the style; let the video own the timing. Choose Matches Video when performance is boss, Matches Image when the photo's facing and a prompted camera matter more. Fix framing mismatch, occlusion, and cuts in the source files before you rewrite adjectives. On Versely, prefer the motion-transfer tool and the named agent job, confirm per-second credits on the live rows, caption and post after the take, and file the still, the cleared driver, and the scene line in a collection so the next character reuses the same scarce asset: a performance you actually have the right to copy.