Guides

    Dolly in or zoom in? Name the mechanism

    A dolly and a zoom both read as getting closer, so models pick at random. How to name the physical mechanism and read the background to check what you got.

    Versely Team9 min read

    Write "the camera moves closer to her" and you have described two different shots at once. One of them puts a camera on wheels and rolls it forward. The other leaves the camera bolted where it is and narrows the lens. They produce different pictures, they mean different things, and a model handed the ambiguous version will pick between them at generation time — often differently on the next take with the same prompt.

    This is one of the most common camera failures in generated video, and it is entirely a naming problem. The fix is to stop describing the effect and start describing the mechanism.

    The two moves are not the same picture

    A dolly in (also called a push-in) moves the camera through space. Because the camera's distance to the subject shrinks proportionally more than its distance to the background, the subject grows faster than the background does. You also get parallax: foreground objects slide across the frame faster than distant ones, and you begin to see around things you could not see around before. Occlusion relationships change. A pillar that hid a doorway stops hiding it.

    A zoom in changes the lens, not the position. The whole image magnifies uniformly. Subject and background grow at exactly the same rate, so their relative sizes never change. No parallax, no new angles on anything, no occlusion changes — a zoom cannot reveal what was hidden, because the camera never moved to a place where it could see it.

    That difference is why the two moves feel different to a viewer. A dolly is a movement through a space, which reads as involvement. A zoom is an act of looking harder from where you already stand, which reads as observation, or surveillance, or a sudden reaction. Neither is the "cinematic" one. They mean different things.

    Name the mechanism, not the effect

    The phrasing that survives is a description of what the physical camera is doing. Compare:

    Ambiguous Unambiguous
    The camera moves closer The camera physically travels forward toward her on a dolly, focal length unchanged
    Zoom in on the product The camera stays fixed in place while the lens zooms from wide to telephoto
    Push in slowly Slow dolly in, the camera tracks forward along the floor toward the subject
    Pull back to reveal the room The camera physically retreats backward, revealing the room around her
    We get tighter on his face The lens narrows to a tight framing on his face; the camera does not move

    Two clauses do most of the work in the right-hand column. The first names the mechanism — travels, retreats, tracks forward for a dolly; the lens narrows or widens for a zoom. The second names what the other mechanism is not doing: "focal length unchanged" for a dolly, "the camera does not move" for a zoom. That second clause is cheap insurance. It closes the door on the interpretation you did not want.

    Keep the whole thing in the first slot of the prompt, ahead of the subject. Camera language placed after the subject and context tends to get approximated or dropped, which is a separate failure from the one this post is about but compounds with it — see camera instructions models silently ignore for why late-position camera language degrades.

    How to tell which one you actually got

    You cannot judge this from a thumbnail. Scrub to the first frame and the last frame of the clip and compare them directly.

    1. Check the subject-to-background ratio. Measure the subject's height against a fixed background feature — a window frame, a horizon line, a doorway. If the subject grew and the background feature grew by the same proportion, it was a zoom. If the subject grew noticeably more than the background feature, the camera moved.
    2. Look for parallax at the frame edges. Play the clip and watch a foreground object against something behind it. In a dolly, the near object slides relative to the far one. In a zoom, they stay locked together and simply scale.
    3. Look for occlusion changes. Does anything become visible that was hidden at the start, or vice versa? Only a dolly can do that. If nothing new appears from behind anything, you got a zoom no matter what the prompt said.
    4. Check the scaling center. A zoom expands uniformly outward from roughly the optical center of the frame. A dolly expands along the direction of travel, which is why the geometry of a dolly looks less like a crop and more like walking.

    The first frame and last frame test is the one to standardize on, because it takes ten seconds and it is unambiguous. The middle of the clip can be misleading — an easing curve that starts slow will look like almost nothing happened at the two-second mark on either move.

    Which one you actually want

    Reach for a dolly when the shot is about entering a space. Push-ins on a face for a realization beat, product hero shots where you want the object to feel like it exists in a room, anything where the audience should feel like they are moving. The perspective change is the point.

    Reach for a zoom when the shot is about the act of noticing. Sudden tightening on a detail, a documentary or surveillance texture, comedic emphasis. Crash zooms in particular are a zoom, not a dolly, and they carry a specific genre signal — the crash zoom and whip pan guide covers why high-velocity versions of these moves have their own failure profile.

    Reach for neither when a cut would do it better. Two static shots at different sizes, cut together, is a completely legitimate way to get closer, and it is the version that never comes back wrong. On a platform where the cut rate is already fast, a viewer will not miss the move.

    The dolly zoom, and why prompts rarely land it

    The Vertigo shot — dolly in while zooming out, so the subject stays the same size and the background field of view warps around it — is the one case where you genuinely need both mechanisms in one prompt, moving in opposite directions at a matched rate.

    It is also compound camera choreography, which is the category where generated video degrades hardest. Two simultaneous camera behaviors have to stay rate-matched for the effect to exist at all; if either drifts, you get a normal-looking push with some background weirdness rather than the effect. Expect a low hit rate from prose alone.

    Two better routes. First, describe it as a transition between two fixed endpoints rather than as a move: first-and-last-frame generation lets you supply the start and end images, and the subject-stays-the-same-size constraint is much easier to express as two frames than as an adjective. Second, use a model that exposes real motion control rather than inferring the move from prose — Versely's model catalog marks which video models carry an explicit camera-control capability, and the motion control comparison narrows the list further.

    The habit worth keeping

    Every time you write a camera line, read it back and ask whether a camera operator could follow it without asking a question. "Move closer" gets a question. "Roll the dolly forward two metres, do not touch the zoom" does not. The prompt that a human crew could execute without clarification is usually the one a model executes without improvising.

    The rest of the movement vocabulary — pans, tilts, tracking, orbits, cranes — has the same property and is covered in camera movement prompts for AI video. Dolly and zoom just happen to be the pair that collide most often, because English gives them the same shorthand.

    FAQ

    Is "push in" a dolly or a zoom?

    Push in is standard shorthand for a dolly in, and pull back for a dolly out. In generated video it is safer to spell it out anyway, because the model is matching text against training captions where the term may have been used loosely. "Push in" plus "the camera physically travels forward" costs you six words and removes the ambiguity.

    Do models handle zooms better than dollies, or the other way around?

    Neither reliably, and the answer varies by model family. A zoom is geometrically simpler — uniform magnification with no new information required — so it is a reasonable prior that zooms come back cleaner. A dolly has to invent parallax and reveal previously hidden geometry, which is a harder generation problem. Test both on the model you actually use before assuming.

    Can I specify how fast the move happens?

    Only loosely. Speed and easing are not inferable from text, so "slow dolly" gets re-estimated on every generation and lands at a different pace each time. If you need a repeatable move across a sequence of shots, that is the point to switch from prose to a model with parameterized camera control, where speed and stop position are set as values rather than adjectives.

    What if I want closer and do not care how?

    Then say so plainly and let the model choose — but say it in one clause, not two. The problem is not ambiguity by itself, it is contradictory precision, where "the camera zooms in as it dollies forward" gives the model two mechanisms and no rate relationship between them. One vague instruction is more stable than two conflicting exact ones.