Fast-motion subjects and the models that survive
Subject speed breaks video models before camera speed does. A per-frame overlap test that predicts limb doubling, plus the shot design that buys real headroom.
Camera moves are a solved-enough problem. Every major video family now takes a named camera movement in the prompt and does something reasonable with it, and a slow dolly-in reliably produces a slow dolly-in. Subject speed is a different question entirely, and it is the one that decides whether a sports clip, a product-reveal snap or a dog shaking off water comes back usable.
The two are not the same load on the model. A camera move translates the whole frame coherently. A fast subject asks the model to move one region a long way while everything else stays put, and to keep that region recognisable as the same object on the other side of the move.
Why subject speed is the harder half
Video models generate frames by predicting what the next one plausibly looks like given the last. That prediction leans heavily on overlap: if a limb occupies roughly the same pixels in frame N and frame N+1, the model has something concrete to carry forward. When the limb moves far enough that its frame-N position and frame-N+1 position do not intersect at all, there is no correspondence to carry, and the model resolves the ambiguity by inventing.
Invention at that moment looks like one of four things, and they show up in a consistent order as speed increases.
- Softening. The fast region loses detail before anything else visibly breaks. Fingers become a mitten, spokes become a disc. This is the earliest tell and the easiest to miss, because it reads as motion blur.
- Stutter. Position jumps rather than travels. The subject is at A, then at C, with no B. On a 200-frame clip this reads as a hitch rather than a defect, and people often blame the export.
- Doubling. The model resolves the correspondence ambiguity by producing both answers. Two forearms, two legs mid-stride, a ball with a ghost of itself trailing behind that is not motion blur but a second ball.
- Detachment. The fast part separates from the body. A hand continues without the wrist, a foot lands unattached. This is the terminal case and it is not recoverable in post.
The order is useful diagnostically. If you are seeing doubling, you were already past the softening threshold two notches back, which means slowing the action slightly will not fix it. You need a structural change to the shot.
The speed ceiling, in frames rather than adjectives
Here is a test you can run on paper before you generate anything. It is not precise, and it does not need to be. It is a triage.
Versely renders at 25 frames per second by default, not 24, so one frame is 40 milliseconds. For any moving feature, compute:
overlap ratio = (distance the feature travels in 40 ms) / (the feature's own width)
Under 1, the feature overlaps itself frame to frame and the model has correspondence to work with. Between 1 and 2 it is marginal. Above 2 you should expect doubling.
Run it on four common briefs, with rough real-world numbers:
| Subject | Speed | Distance per frame | Feature width | Ratio |
|---|---|---|---|---|
| Hand snapping fingers | ~2.5 m/s at the fingertip | ~10 cm | ~9 cm finger arc | ~1.1 |
| Dog shaking its head | fast rotation, small radius | small linear travel | whole head | under 1 |
| Sprinter at full speed | ~10 m/s | ~40 cm | ~15 cm thigh | ~2.7 |
| Thrown ball | ~20 m/s | ~80 cm | ~20 cm ball | ~4.0 |
That table matches what people actually report. Snaps are borderline and sometimes work. A dog shaking is mostly rotation of a large object, so despite feeling fast it is a comparatively easy ask. Sprinting reliably produces leg doubling. Ball flight rarely survives at all, which is why generated sports footage almost always cuts away from the ball.
Two important caveats. The ratio is about the feature, not the subject, so a walking person swinging their arm quickly can fail at the hand while the body is fine. And it assumes the frame is fixed, which is precisely the assumption the best workaround breaks.
Shot design that buys headroom
Every one of these attacks a term in the ratio.
Track the subject. This is the highest-leverage move available and it is worth internalising. If the camera moves with a sprinter at the same speed, the sprinter's displacement in frame space drops close to zero. The legs still cycle, but the body no longer crosses the frame. You have converted a subject-speed problem, which models are bad at, into a camera-move problem, which they are good at. Write it explicitly: "tracking shot moving with her at matched speed, subject held centre frame." Camera control is doing the work here, not the motion setting.
Frame wider. Per-frame displacement matters as a fraction of frame width, not in centimetres. The same sprinter filling the frame moves a large fraction of it per frame; the same sprinter as a small figure in a wide establishing shot moves a small fraction. Wide plus fast is survivable. Tight plus fast is not.
Cut around the peak. Almost every fast action has a slow beginning and a slow end. The wind-up before the throw and the follow-through after it are both low-ratio. The 200 milliseconds in the middle are the only part that breaks, and they are also the part the audience reconstructs on its own. Generating the wind-up and the aftermath as two clips and cutting between them is standard editorial practice, not a compromise.
Shoot the consequence. The net moving, the water already on the floor, the sprinter's chest crossing the line. These carry the same story beat at a fraction of the ratio.
Raise the frame rate where the model offers it. A subset of the catalog exposes a frame-rate option that includes 60fps, among them Veo 3.1, Sora 2 text-to-video and Kling 2.5 Turbo. Generating at 60 rather than 25 cuts the per-frame travel to roughly 40 percent of what it was, which moves a 2.7 down near 1.1. This is the only lever on the list that changes the model's job directly rather than changing your shot, and it is underused. See frame rate for how the setting interacts with delivery.
Use the motion parameter, not adjectives. Where a model exposes motion level as Low, Medium or High, that is the actual control. Writing "extremely fast, explosive, high-energy" into the prompt raises the register of the scene without changing what the model was going to do with the limbs.
What not to try
Slowing it down afterwards. Changing playback speed redistributes existing frames. If frame 40 contains two left legs, slowing the clip gives you two left legs for longer.
Frame interpolation as a rescue. Interpolation synthesises intermediate frames from neighbours. Where the neighbours already disagree about how many limbs exist, it interpolates between two wrong answers and produces a smoother wrong answer. It is a good finishing tool on clean footage and a poor repair tool on broken footage.
Prompting harder. There is no phrasing that gives the model correspondence it does not have. This is a geometry constraint, not a comprehension one.
Testing your own ceiling
Do this once per model you care about and keep the answer.
- Write one action with a stated speed, not an adjective. "A hand snaps its fingers once, close on the hand" is testable. "Dynamic hand movement" is not.
- Generate the same action at three framings: tight, medium, wide. This isolates the framing term of the ratio.
- Generate the winner at three seeds. One good sprint out of three is not a model that does sprints.
- Review frame by frame, not in playback. Doubling at 25 frames per second can pass unnoticed at speed and be glaring on a pause, which is exactly what a viewer scrubbing back will do.
- Record the framing at which the action first survived. That is your ceiling for that model, and it will be stable until the model version changes.
The agent chat can fan one prompt across several named models in a single request, which makes step 2 a few minutes rather than an afternoon. If you want a second opinion on a borderline clip, asking for feedback on a generation will flag artefacts you have stopped seeing after the twentieth viewing.
FAQ
Which model handles fast motion best?
Framing and frame rate move the outcome more than model choice does, and any ranking published today has a short shelf life. What generalises: models exposing a 60fps option are structurally better placed for fast subjects, and image-conditioned paths start from a fixed subject rather than an invented one. Test your specific action rather than trusting a leaderboard position, since arena scores reward general preference, not limb integrity at speed.
Why does slow motion look good but real-time speed look wrong?
Because a slow-motion brief is a low-ratio brief. You are asking for small per-frame displacement, which is the regime where the model has correspondence to work with. That is genuinely why so much generated sports and action content is in slow motion, and it is a reasonable choice rather than a cheat.
My clip is fine until the last second, then it falls apart. Why?
Drift accumulates. Each frame is predicted from a state that is already slightly wrong, and a fast subject amplifies the error faster than a slow one. Shorten the duration and generate two clips instead of one; the join will cost you less than the collapse does.
Does this apply to hands specifically?
Hands are the worst case because they combine a high ratio with heavy self-occlusion, which is covered in more detail in our piece on the artifacts that still expose AI footage. A moving hand is failing two tests at once, which is why a static hand holding a product works and the same hand performing a quick gesture does not.