Where motion coherence breaks as clips get longer
A 15-second ceiling is rarely 15 usable seconds. The three ways long clips fall apart, a ladder test to find your safe duration, and where to place the cut.
Forty-three video models in the catalog top out at 15 seconds and forty-one top out at 10. Those are the two big clusters, and both numbers are read as if they were durations you can use. They are dispatch caps. The number of seconds that survive review is a different, smaller number, and it belongs to the shot rather than the model.
The useful skill is not finding the model with the highest ceiling. It is knowing which of three specific failures your shot will hit first, roughly when, and where to put the cut so the audience never sees it.
Three ways coherence breaks
Drift is not one phenomenon. It arrives as one of three, and which one you get depends on what is in frame.
Subject morph. The subject slowly stops being the same subject. Face geometry shifts, a jacket changes its collar, hair length creeps, a logo on a shirt reorganises itself. This is the most common failure on any shot with a person and the hardest to notice in real time, because each frame is only slightly different from the last. The reliable check is to compare the first and last frame side by side rather than watching the clip through.
Background reshuffle. Objects behind the subject rearrange. A doorway relocates, a lamp appears, the number of chairs at a table changes, a wall's texture becomes a different wall. Busy backgrounds fail earlier than plain ones, and backgrounds with countable objects fail most visibly because the audience can count.
Gait reset. Cyclic motion loses its phase. A walking subject takes a step that does not follow from the previous one, an arm swing inverts, a rotation reverses direction mid-turn. This is the most jarring of the three for a viewer, because human motion perception is unforgiving about legs. It also arrives earliest on shots where the subject moves through the frame rather than in place.
Two useful patterns. First, plain backgrounds buy you seconds; busy ones cost them. Second, motion complexity dominates everything else, so a static portrait at 15 seconds is a much safer bet than a walking subject at 8.
Why the ceiling and the usable window are different numbers
The ceiling is set by what the model was trained and served to produce in one pass. The usable window is set by how long the model can hold your particular scene's state before its own accumulated output starts to dominate its conditioning. Longer clips give the model more of its own earlier frames to condition on, and any small error is now part of the input.
Duration ladders make this concrete. Grok Imagine offers 6, 10, 15, 20 and 30 seconds, and Sora 2's storyboard mode offers 10, 15 and 25, but neither ladder tells you which rung holds. Meanwhile Pruna P-Video exposes every whole second from 1 to 20 and Kling V3 and Happy Horse expose every second from 3 to 15, which matters enormously once you know your safe number: a ladder with whole-second granularity lets you buy exactly the seconds that work. A model offering only 5, 10 and 15 forces you to buy 15 and trim, or buy 10 and lose the end of the beat.
There is also a structural split worth knowing. Some long durations come from a genuine single pass and some come from continuation, where the model extends its own output. FLUX.3's extend entry and the LTX extend entries are explicit about this, and it changes the failure profile: single-pass long clips drift gradually, while continuation-based clips tend to hold well and then show a visible seam at the join. Different problem, different fix, and how far extension actually gets you covers the practical limits.
The ladder test
One session per model and shot type, and the output is a number you keep.
- Pick the hardest shot you actually need. Not a test pattern. If your campaign has a walking subject in a busy interior, test that, because a locked-off portrait will pass at every duration and teach you nothing.
- Generate at three rungs: roughly a third of the ceiling, roughly two thirds, and the ceiling itself. Fix the prompt and the seed across all three.
- Score each clip on the three failures separately. Note the timestamp of first subject morph, first background reshuffle, first gait reset. Watching for one thing at a time is much more reliable than watching for "problems."
- Take the earliest of the three timestamps across the runs. That is your drift point for this shot type on this model.
- Subtract about a second. Cut before the drift, not at it, because the frames immediately preceding a visible failure are usually already degraded.
That last number is the one to write down. It is stable for a given combination of model, shot type and background complexity, and it stops being valid when the model version changes.
Where to place the first cut
| Class | Declared ceiling | Starting assumption for the first cut | What usually breaks first |
|---|---|---|---|
| Locked-off subject, plain background | Whatever the model offers | Close to the ceiling | Background reshuffle, late |
| Talking head, mid shot | 10–15s | Around two thirds of the ceiling | Subject morph on wardrobe and hair |
| Product on turntable or slow orbit | 10–15s | Around two thirds | Background reshuffle, then label detail |
| Person walking through frame | 10–15s | Around half | Gait reset, early |
| Busy interior or street, any subject | 10–20s | Around half | Background reshuffle |
| Crowd or multiple subjects | Any | Around a third | All three, in sequence |
| Continuation or extend-based | 20s+ | Cut on the join | Seam at the extension boundary |
Those fractions are placeholders, not measurements. They are where to point the ladder test first so you are not sampling blind. Replace each one with your own number the first time you run the test on that shot type, and the table stops being generic advice and becomes your production spec.
Cutting before the drift instead of after
Once you have a safe duration, the fix is editorial rather than technical. Long AI clips do not need to be long clips.
- Generate to your safe duration, not the ceiling. Buying seconds you will trim is the most common avoidable spend in this whole workflow.
- Cut on motion. A cut placed on a fast movement, a turn, or a gesture hides the fact that the next shot is a separate generation. A cut on a still frame invites comparison.
- Change the framing across the cut. Wide to mid, mid to close. Two clips at the same framing read as one clip that jumped; two clips at different framings read as coverage.
- Let the background change. Trying to hold an identical background across a cut is the hardest thing to ask of a generator, and the audience does not require it if the framing changed.
- Reserve the long take for shots that earn it. A single unbroken 15 seconds is a specific creative statement. If the shot does not need it, three cuts are easier, cheaper and more robust.
When you genuinely need continuity across a boundary, the two working approaches are keyframe chaining, where the last frame of one clip seeds the next, and continuation prompting, where you feed the model its own final seconds. Both trade a little quality at the join for continuity you control, which is usually the better deal than gambling on a single long pass.
If drift is showing up between separate shots rather than inside one, that is a different diagnosis with different causes, and consistency collapse between shots is the right starting point.
FAQ
Is a model with a 30-second ceiling better for long content?
Only if the shot survives 30 seconds, and most do not. A high ceiling is genuinely useful for a locked-off shot with a plain background where nothing complex is moving. For a walking subject in a busy environment, a 30-second ceiling mostly means you can spend more to generate footage you will cut anyway. Test the shot before treating the ceiling as an advantage, and read the state of the duration race for what the long-ceiling tier is actually shipping.
Does higher resolution make drift better or worse?
Neither reliably, but it makes drift more visible. The same subject morph that reads as softness at 720p reads as a different face at 4K. If you are moving a proven shot up a resolution tier, re-check the drift point rather than assuming it carries, because the artifact threshold moved even if the model's behaviour did not.
Can I prompt my way out of drift?
Partly. Explicit continuity language helps: naming the wardrobe, naming the background elements that must persist, specifying a single continuous camera move rather than leaving the camera to the model. What prompting cannot do is extend the window indefinitely, and past a certain duration more prompt detail starts to hurt, because you have given the model more specific things it can contradict. Treat prompting as worth one or two seconds, not as a fix.
What is the fastest way to find my safe duration without burning a lot of credits?
Test one shot type on one model at the ceiling only, and read the timestamps. A single clip at the maximum tells you where all three failures land; the lower rungs mostly confirm what the top rung already showed. Then run the agent with the same prompt across a couple of candidate models in one request so the comparison is clean. Two clips and ten minutes of careful watching is usually the whole exercise.
The duration index shows which models even offer the seconds you need. Everything after that is your ladder test, and the number it produces belongs on the shot list next to the model name.