Depth of field and rack focus in video prompts
Shallow and deep focus are executable prompt instructions. A focus pull is not, unless you stage it across two generations. How to specify the focus plane.
Ask a video model for a shallow depth of field and you will usually get one. Ask it for a rack focus from the coffee cup to the barista behind it and you will usually get a clip where the focus never moves, or where the whole frame softens and re-sharpens for no reason. The difference between those two outcomes is not a phrasing problem. It is a category problem.
Depth of field is a state. The model can hold it across every frame of a clip because it is a property of the picture, no different from the grade or the lighting. A focus pull is an event — a change of state that has to happen at a particular moment and complete in a particular time. Camera events of that kind are the weakest thing you can put in a prompt, and focus racking is at the bottom of even that list, because most models have no explicit concept of a lens changing its focus plane mid-shot.
Once you sort your request into state or event, what to do about it becomes obvious.
Specifying the focus plane, not the vibe
The failure mode with depth of field is writing it as an adjective. "Shallow depth of field" on its own tells the model to blur something and leaves it to pick what. Sometimes it picks the background. Sometimes it softens the subject's ears and leaves the wall sharp.
The fix is to name both sides of the plane. Say what is sharp and say what is soft, and give each one a concrete object rather than a zone:
Shallow depth of field. Sharp focus on the espresso cup in the foreground; the barista and the machine behind her fall soft and out of focus.
Deep focus. The keys in the foreground, the desk in the midground and the window at the back of the room are all sharp.
Three things are doing work there. The state is named (shallow depth of field, deep focus). The sharp side is anchored to a specific object. The soft side is anchored to a different specific object. A model that has all three has a plane to place. A model that has only the first is guessing.
Aperture notation helps as a secondary cue. f/1.4 correlates with shallow and f/16 with deep in the captions these models learned from, so it nudges in the right direction, but it does not name a plane on its own. Write it alongside the plain-English version rather than in place of it.
Keep all of this in the cinematography slot at the front of your prompt, next to shot size, angle and focal length. Depth-of-field language that arrives after the subject and lighting is competing with everything already established, and late camera language is exactly what gets approximated away — the pattern is documented in camera instructions models silently ignore.
If you actually need the focus to change during the shot, stop asking the lens to do it and restructure the shot instead. There are three routes, in ascending order of cost and reliability.
Route one: move the subject through the plane
This is the cheapest option and it works on models that have no focus-pull behaviour at all. Instead of racking focus onto a person standing in the background, keep the focus plane fixed and have the person walk into it.
Shallow depth of field, focus fixed on the doorway in the midground. A man walks out of the soft background and into the sharp doorway plane, becoming crisp as he arrives.
What you have written is a subject-motion instruction, which models handle far better than a camera event. The audience reads the result as a focus change. Mechanically it is not one, and that is the point.
Route two: pin both ends with first-and-last-frame
First-and-last-frame generation lets you supply the start image and the end image and have the model fill the middle. That is the natural shape for a rack focus, because a focus pull is exactly a transition between two clearly definable states.
The workflow:
- Generate a still with the foreground sharp and the background soft.
- Generate the matching still with the background sharp and the foreground soft, holding the framing identical. Same framing matters more than anything else here — reuse the same base image and edit the focus rather than generating two independent stills, or the interpolation will have to invent a camera move to reconcile them.
- Feed both as the endpoints. Describe the transition in the prompt as a focus change and nothing else:
the focus plane travels from the foreground cup to the woman behind it. The camera does not move.
Models that accept both endpoints include Veo's first-last-frame variant and Flux 3's first-last-frame model. The broader technique, including how to chain several of these into a longer sequence, is covered in the first-last-frame film workflow.
Route three: generate two clips and dissolve between them
The blunt version, and the most reliable. Generate the same shot twice with the same framing, one focused near and one focused far, then cross-dissolve from the first into the second in the timeline.
The dissolve length is what sells it. A real focus pull on a fast lens takes roughly a third of a second to a second depending on how far the plane travels. At Versely's default frame rate of 25 fps, that is a dissolve somewhere between 8 and 25 frames. Shorter than 8 frames reads as a cut; longer than about 25 frames stops reading as a focus pull and starts reading as a dissolve between two shots, which is a different edit entirely.
The video editor handles this as a normal crossfade transition between two clips on one timeline, and because the timeline is EDL-based rather than baked, you can adjust the dissolve length and re-render without regenerating anything. Checking the timing costs nothing: preview: true gives a free 480p pass, subject to a short per-user cooldown, so you can watch the dissolve at three different lengths before committing. The final export is charged once regardless of how many clips are on the timeline, which is what makes the two-clip approach practical rather than wasteful — the preview and export cost breakdown has the details.
What breaks a staged focus pull
Framing drift between the two clips. If the second generation sits two degrees off the first, the dissolve turns into a small unmotivated camera jump. Generate both from the same source image where possible, and if you cannot, at least hold the seed constant and change only the focus language.
Bokeh that wobbles. Some models produce a background blur that pulses frame to frame rather than staying constant. It reads as a technical artifact, not a lens. Shortening the clip usually helps, because the wobble compounds over length. If it persists on one model, it is worth trying the same prompt on another before rewriting it.
The model reverting to deep focus. A shallow depth of field stated once at the start of a long generation often decays by the end, with the background quietly sharpening up. Same fix as most drift problems: shorter clips, or pin the end state with a first-and-last-frame generation so the tail has a target.
The subject leaving the plane. If you have specified a shallow plane and then asked the subject to walk toward camera, they will pass out of focus, correctly. That is either what you wanted or a contradiction you did not notice. A shallow plane and a subject moving in depth need each other's permission.
FAQ
Does "bokeh" work as a prompt word?
As a nudge, yes. It reliably biases toward out-of-focus highlights, particularly the round or oval blobs from point light sources. It is not a substitute for naming the plane, because it says nothing about what is out of focus. Use it as a texture note in the style slot, with the actual sharp-and-soft specification in the camera slot.
Can I get a shallow depth of field on a wide-angle shot?
Less easily, and that is physically correct. Short focal lengths produce deeper apparent focus, so a prompt asking for 24mm wide-angle and very shallow depth of field in the same line is asking for two things that fight each other. If you need heavy background separation, move to a longer lens in the prompt and let the framing follow.
Why does my focus pull happen but at the wrong moment?
Because timing is not inferable from prose. "Focus pulls to her at the halfway point" gets re-estimated on every generation and lands somewhere different each time, the same way "slow dolly" does. If the beat matters, stage the pull as a dissolve between two clips in the editor, where the frame it happens on is a value you set rather than a phrase you hope lands.
Is a focus pull worth the extra work at all?
On a hero shot or an ad beat, often yes, because it directs attention in a way a cut cannot. On general b-roll, usually not. Two static shots at different focus depths, cut together, communicate almost the same thing for a fraction of the effort, and they never come back wrong.