There is no camera in the scene, so this is really instruction about how the frame should change over time. Models expose it in two ways: as prompt vocabulary — dolly in, orbit left, handheld — or as an explicit setting, most commonly a switch that locks the camera off entirely.
The locked-off switch earns its keep more often than the exotic moves. Generated camera drift is the most common reason a clip looks synthetic: nobody asked for movement, the model added a slow float anyway, and the shot reads as unstable. Fixing the camera also gives the model one less thing to spend its coherence budget on, which usually improves the subject.
Camera language and subject language compete. Ask for a fast orbit around a person who is also running and something has to give, usually the person's anatomy at the fastest part of the move.
In practice
- Name one move, not three. "Slow push in" is a shot; "push in while panning and craning" is a mess.
- A locked camera is the right default for product shots, talking heads and anything with text on screen.
- Camera terms are style keywords, not physics — the model approximates the look of the move.
The mistake to avoid
Describing camera work in an image-to-video prompt while also re-describing the frame. The model tries to satisfy both and drifts the composition away from your input still.
Where you will run into it
- NYC Street Interview — Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
Related terms
Motion level
Motion level is a coarse dial for how much movement a generated clip contains, usually offered as low, medium or high rather than a precise number.
Style preset
A style preset is a named look — cinematic, anime, documentary — that a model applies without you having to describe it in the prompt.
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
Seed
A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.