Physics Failures: Weight, Gravity and Contact in Generated Motion
Video models predict plausible next frames, not physical forces. The VideoPhy numbers, four failure categories, and what actually reduces them in practice.
A prompt I run often as a sanity check: a stack of three cardboard boxes, the top one shoved off the edge of a table. In a good take, the box tips at the pivot point, hangs for a beat past its center of mass, then drops and lands with a compression you can feel. In most takes, something smaller is wrong — the box doesn't accelerate as it falls, or it lands and just stops with no give, or it phases half an inch into the table before settling. Nobody watching would call it "broken." They'd just feel that something about the weight was off, and they'd be right.
That's the shape of physics failure in generated video: rarely a glitch you can point at, usually a plausibility gap between what a camera would have recorded and what a model decided was close enough. It's also, per the one benchmark that's actually tried to measure it, the single biggest gap between "looks impressive" and "is correct" in the category.
Why physics breaks before anything else
A video model isn't running a physics engine. It's trained to predict a plausible next frame — or, in modern architectures, a plausible next block of latent frames — from statistical patterns in footage it's seen. That's a fundamentally different task from simulating forces. A physics engine starts from mass, velocity, and gravity and integrates forward; a video model starts from "what usually follows a frame that looks like this" and renders forward. Most of the time those two answers land close enough that the difference doesn't register. Motion that depends on invisible quantities — how much something weighs, how much force a hand is applying, what happens at the exact instant two surfaces touch — is where they diverge.
This is a different failure from temporal consistency, which is about a shirt staying the same color or a background staying put from frame to frame. A clip can be perfectly temporally consistent and still get the physics wrong — the falling box example above has no flicker, no drift, nothing an eye would flag as an artifact in a paused frame. Play it, though, and the motion reads as smooth in the way a video game's canned animation reads as smooth: plausible-looking, not force-derived. Consistency asks "did it stay the same." Physics asks a harder question: "did it move the way mass and momentum actually dictate."
What the numbers actually say
Until recently, "does this obey physics" was a subjective call — something a reviewer felt rather than something anyone measured. VideoPhy is the first benchmark built specifically to score it, testing generated video against real-world physical rules across three interaction types: solid-solid (stacking, collisions, one object striking another), solid-fluid (pouring, splashing, an object dropped into liquid), and fluid-fluid (two liquids mixing or separating).
The result that matters: across the models tested, the strongest performer — CogVideoX-5B, an open research model with no relationship to anything in Versely's catalog — adhered to both the input caption and physical commonsense on only 39.6% of instances. Read that plainly: the best model in the study got the request right and obeyed physics together in roughly two out of five outputs. The rest either drifted from what was asked, broke a physical rule, or both. The same paper introduced an automated evaluator, VideoCon-Physics, specifically because scoring this by hand didn't scale — which is itself a tell about how early this axis of quality still is compared to resolution or prompt adherence, both of which have had measurable benchmarks for years.
Four ways it actually shows up
Watching enough generations fail this way sorts into a short list of recognizable patterns:
- Weight and mass. A boulder and a beach ball fall at visually similar rates; a performer lifts a "heavy" prop with the same ease as an empty box. Nothing in the model's training told it these two objects should move differently, because nothing about mass was ever labeled — it only ever saw pixels.
- Gravity and free-fall. Objects drift downward at something closer to constant velocity than acceleration, or hang for an extra beat before "deciding" to fall — a tell that the model is choosing plausible-looking positions frame to frame rather than integrating a force over time.
- Contact and collision. A hand closes around a cup with no deformation at the grip point, or passes slightly through it. A struck object doesn't transfer momentum to what it hits. Two surfaces meet without either one absorbing the impact.
- Fluid and soft-body behavior. Poured liquid doesn't fill a container correctly, or a cloth or cable holds a shape with no tension running through it — soft materials rendered with the internal logic of a rigid one.
Solid-solid contact is the most forgiving of the three VideoPhy categories precisely because it's the one closest to what a model has seen the most of; fluid interactions are the least forgiving because the internal degrees of freedom are highest and the training signal for "correct" fluid behavior is sparser.
Where it actually costs you
Most brand and social work never triggers this at all — a talking-head clip or a slow push-in on a product has no physics to get wrong. It shows up specifically where the brief includes a physical event: an unboxing where the lid doesn't seat, a product demo where a cap doesn't seal, a food clip where a pour looks weightless, any B-roll where two things are supposed to touch convincingly. Those are also, not coincidentally, some of the highest-intent shots in commerce content — the moment a viewer is meant to believe the product is real and works the way it's shown working. A physics tell in exactly that moment undercuts the one thing the shot needed to do.
What actually reduces it
None of this is fixable by asking the model to "obey gravity" in the prompt — mass and force aren't concepts a diffusion or transformer video model reasons about, so there's no instruction that makes it start integrating forces. What does move the needle:
- Shorten the window around the physical event. Fewer seconds means fewer frames where the model has to keep integrating an invisible force correctly. A two-second contact shot has far less room to drift than an eight-second one.
- Cut around the contact instead of through it. Show the box tipping, cut, show it settled on the floor. The moment of impact is exactly where models are weakest, and it's also the one moment an editor can skip without the audience noticing.
- Keep the camera still. Camera motion adds degrees of freedom the model has to resolve simultaneously with the object's motion; a locked-off shot isolates the physics problem instead of compounding it.
- Write the physical detail into the prompt anyway. It won't make the model simulate force, but naming weight, speed, and material ("a heavy glass jar, falls hard, cracks rather than shatters") biases generation toward footage where those cues were more often correct in training. This is the same lever covered under prompt adherence generally — it raises the odds, it doesn't guarantee the outcome.
- Bound the shot with a start and end frame. First-last frame generation fixes the two moments a viewer actually judges — before and after — and lets the model fill the physically uncertain middle in between, rather than leaving the whole arc to chance.
- Test the specific shot across models before you commit to one. Physics adherence isn't evenly distributed across the field, and which model handles a stack of boxes convincingly isn't the same one that handles a pour convincingly.
Testing it directly in Versely
That last point is worth doing as an actual step rather than a guess. Versely's agent can dispatch one prompt to several video models in a single request and hand back the results side by side, which turns "which model handles this contact moment better" from a guess into a five-minute comparison. A prompt like:
"Generate this on VEO 3.1, Kling O3 Pro, and Seedance 2.0 so I can compare the physics: a glass jar tips off a counter edge, falls, and cracks on the tile floor. Static camera, no motion."
runs the same brief across all three in parallel instead of three separate sessions of trial and error. Play the three results back to back on the exact contact frame — does the jar accelerate believably, does the crack line up with impact, does anything interpenetrate — and you have a direct answer for this shot instead of a general reputation for "good physics" that may not hold for your specific brief.
For a standing view of how models stack up beyond a one-off test, the model rankings track quality across the catalog, compare puts two models head-to-head on the same criteria, and the best text-to-video model list is worth checking before a physics-heavy brief specifically, since it's rescored as new releases land rather than fixed at launch.
FAQ
Why does generated video get physics wrong more than other things?
Because video models are trained to predict plausible next frames from pixel patterns, not to simulate mass, force, or momentum. Appearance and motion can look locally smooth while still being physically wrong — the model was never solving for weight or gravity in the first place, only for what usually comes next visually.
Is this the same problem as flickering or inconsistency?
No. Flicker and drift are about an object failing to stay the same from frame to frame. Physics failures can happen in a perfectly stable, non-flickering clip — the object stays consistent, it just doesn't fall, land, or collide the way its apparent mass would demand.
Can better prompting fix physics errors?
It helps at the margins — naming weight, speed, and material biases the model toward training examples where those cues were usually correct — but it doesn't make the model simulate force. Shortening the shot, cutting around the contact moment, and testing across models move the needle more reliably than prompt wording alone.
Which interaction type is hardest for current models?
Fluid-on-fluid interactions score lowest in the VideoPhy benchmark's three categories, followed by solid-fluid. Solid-on-solid contact — objects stacking or colliding — is the most forgiving, likely because it's the pattern models have seen the most of in training.