Describe the camera or hand it a reference clip
Which camera moves survive as text, which get silently approximated, and what a motion-reference clip costs you in framing freedom. With a three-bucket rule.
There are two ways to tell a video model how the camera should behave. You can write it — "slow push in, locked on her face" — or you can hand it a clip and say do this. Almost everybody writes it, because the prompt box is already open and attaching a clip means finding or shooting one first. That's a reasonable default and it's wrong for a specific, identifiable class of move. The useful question isn't which method is better; it's which of your shots belong in which bucket, and there are only three buckets.
Text is a suggestion; a reference clip is a constraint
A prompt is read, parsed and approximated. That's the mechanism, not a failure of it — the model resolves your words into a plausible motion, and "plausible" is doing all the work. Camera instructions models silently ignore goes into how that approximation shows up on screen; the summary is that camera directives rarely error, they just come back softer, later, or shaped differently than what you asked for.
The catalog tells you how much of this is even nominally supported. Of the 146 video models in Versely's catalog, 15 declare a camera_control capability — a minority, and it includes the endpoints you'd expect: Runway Gen-4.5, described by its own listing as having precise camera controls; Seedance 2.0, which claims director-level camera and lighting control; and the Kling O3 Pro text-to-video line. Everything else is taking your camera language as ordinary prose. That distinction is worth checking rather than assuming — get_model_input_schema returns the exact fields, allowed values and defaults a given endpoint accepts, per provider route, which is faster than testing it empirically.
A reference clip works differently. The motion-control family doesn't interpret motion, it transfers it. Versely's catalog carries seven motion-control endpoints, and their own descriptions are blunt about the mechanic: Kling Video V3 Pro Motion Control "transfers motion from a driving video to a reference image," and DreamActor V2 transfers facial expressions and head movements from a driving video to a reference image. Wan 2.2 14B Move sits in the same category. Every one of these requires a video input — exactly one — and it goes in a dedicated field, not in the prompt text. The character image goes in image_url and the driving clip in video_url; embedding a URL in the prompt does nothing.
What the reference clip takes from you
The trade is real and it is under-discussed. A driving clip doesn't just supply a motion curve. It supplies:
- Shot size. If the reference is a medium two-shot, you are making a medium two-shot. You cannot independently ask for a close-up.
- Subject position over time. Where the subject sits in frame at second 0 and second 3 is decided by the reference, not by you.
- Duration envelope. The move takes as long as it took in the clip.
- Timing of every beat. Which is the entire reason you'd use one.
So the honest framing is: text gives you framing freedom and costs you fidelity; a reference clip gives you fidelity and costs you framing freedom. Neither is a strictly better tool, and the mistake is bringing a reference clip to a shot whose framing was the point.
| Written camera direction | Motion-reference clip | |
|---|---|---|
| Precision of the move | Approximate; drifts off-axis, arrives early or late | High — it's transferred, not inferred |
| Framing freedom | Full | Inherited from the reference |
| Duration control | You set it | The reference sets it |
| Setup cost | None | You need a clip that already does the move |
| Model availability | Every video model takes prose | Seven motion-control endpoints, plus reference-video inputs on a handful of others |
| Fails as | A generic version of your move | A perfect move in the wrong frame |
There's a middle option worth knowing about. Some non-motion-control endpoints accept reference videos alongside the prompt rather than as a driving track: generate_videos exposes a reference_video_urls parameter, and in the catalog Wan 2.7 Reference to Video accepts up to five reference images and/or videos with optional first-frame control, while Seedance 2.0 takes up to two reference videos as an optional input. Those are softer than a driving clip — closer to "look like this" than "do exactly this" — and they preserve more framing freedom at the cost of less fidelity.
The three buckets
Always describe. Moves that are a single continuous vector with no timing requirement. Push in, pull out, pan left, tilt up, slow orbit, static lock-off. These are the moves where "approximately correct" is genuinely correct, because there's no beat the move has to hit. Describing them also leaves you free to pick the framing, which is worth more than the last 10% of motion accuracy. Practical wording: name one move, one direction, one speed, and stop. Stacking three camera verbs in one prompt is the most reliable way to get none of them.
Always reference. Moves whose entire value is the timing or the texture. Human performance — a specific gesture, a head turn on a word, an expression change. Handheld character, where the point is the irregularity of a real operator's hands. A whip pan that has to land on a beat. Anything you could describe in words but only badly, because the description would be "you know, like that." These are motion-control jobs, and AI Motion Transfer is the tool surface for them. The best motion control models ranks the endpoints.
Split into two shots. Moves that change both the camera and the subject's blocking mid-shot. "Dolly in as she turns, then crane up over the roof as she walks out of frame." That's two intentions in one instruction, and it's the shape that fails hardest under both methods — text approximates it into a mush, and no reference clip you have will contain both halves in the right order. Cut it. Shot A is the dolly on her turn; shot B is the crane. You get two clean generations, each with a single vector, and a cut you control. That the model can sometimes do it in one pass is not a reason to ask it to — single-pass multi-shot directing covers where that ceiling actually sits.
A five-minute test for which bucket a shot is in
Before committing either way, run this. It costs two generations and settles the argument.
- Write the move as plainly as you can — one verb, one direction, one speed. Generate it.
- Watch it looking only at the camera, not the content. Ask: did it start when I expected, travel where I expected, and stop where I expected?
- If two of those three are right, you're in the describe bucket. Ship it.
- If the move is right but the timing is wrong — it arrives too fast, or lands in the wrong place relative to the action — you're in the reference bucket. The model understood the move and can't hit your beat.
- If the model did something recognisably different from what you asked, or did the first half and abandoned the second, you're in the split bucket. Cut the shot in two and test each half.
Step 4 is the one people misread. A move that's correct in shape and wrong in timing looks like a prompting failure, so the instinct is to add more adverbs. Adverbs don't buy timing. A reference clip does, and it's the only thing that does.
FAQ
Can I use a reference clip for the camera and still control the framing?
Not with a driving clip on a motion-control endpoint — the reference's framing is the output's framing, and that's the deal you're making. The softer reference-video inputs on endpoints like Wan 2.7 Reference to Video leave you more framing latitude, because the clip is being treated as guidance rather than as a track to follow. If framing is non-negotiable and timing is too, that's a signal the shot should be split.
Does the reference clip need to be the same length as my target?
The move takes as long as it took in the reference, so a three-second gesture won't fill a ten-second clip. Cut the reference down to the exact motion you want before submitting it rather than expecting the model to pick the interesting part out of a longer take.
Do camera keywords do anything on models without a camera-control capability?
They're read as prose and factored into the generation like any other words, so they're not inert — but there is no parameter being set and no guarantee being honoured. On those endpoints, treat camera language as influence rather than instruction, and don't spend rerolls trying to force a move the model was never going to execute precisely.
Where does prompting a "handheld" look fit?
Reference bucket, almost always. Handheld is a texture rather than a move, and every model has a house version of it that reads as a gentle sway. If the brief needs the specific instability of a real operator — the tiny overcorrections, the breath — that's transferred from a clip, not described. Camera control covers what the parameter-level version of this can and can't reach.