Guides

    Foot sliding and walk cycles in AI video

    Feet skate because ground contact is never specified. Prompt phrasing for footfall, weight and surface, plus the shot sizes where walking actually holds up.

    Versely Team9 min read

    Generate a character walking and watch the feet. More often than not the legs are doing a plausible walking motion while the body translates across the ground at a speed that has nothing to do with the stride — feet skating, or planting and then sliding a few centimetres before lifting. The rest of the shot can be excellent and this one thing makes it read as animation rather than footage.

    The reason is specific and so is the fix. Walking is the one common human action where the important information is not in the body at all. It is in the contact between a foot and the floor, and nothing in an ordinary prompt describes that contact.

    Contact is the thing nobody specifies

    A video model matches appearance across frames; it does not run a physics simulation. There is no ground plane, no friction coefficient, no notion that a planted foot is a fixed point the rest of the body rotates around. Every frame re-derives where the foot is from what the previous frames looked like, and "looked like" is a much weaker constraint than "was physically attached to." That gap is the general case in physics failures: weight, gravity and contact; walking is the version you hit most often.

    That produces two related failures. Sliding is translation speed disagreeing with stride length — legs cycling at one rate, ground passing at another. Gait looping wrong is the cycle itself breaking: a stride that never completes, a leg that swaps which one is leading, a step that resets rather than continues.

    Both come from the same gap. The prompt described a person walking. It did not describe a foot landing, taking weight, and staying put while the body moves over it. Speed is not inferable from text either, which is why "slow walk" gets re-estimated on every generation.

    Phrasing that gives contact something to key on

    Three properties do the work: footfall, weight, and surface. Say all three.

    Vague Specific
    "walking down the street" "heel strikes first, then the ball of the foot; each foot stays planted as the body passes over it"
    "walking slowly" "one deliberate step per beat, weight settling fully onto the front foot before the back foot lifts"
    "she walks toward the camera" "she takes four measured steps toward the camera, shoulders level, arms swinging in opposition to the legs"
    "on the ground" "on wet asphalt, shallow puddles, shoes leaving brief dark prints"
    "walking confidently" "long stride, heel landing hard, weight forward over the toes"

    Two things are happening in that right-hand column. First, the description names a mechanism rather than an effect. "Walks confidently" is an effect; "heel landing hard, weight forward" is something the model can render.

    Second, it makes contact explicit as a constraint: the foot stays planted while the body passes over it. That clause does more work than any adjective in the prompt, because it is the only part describing the thing that actually breaks.

    Countable steps help too. "Four measured steps" is a schedulable event count; "walking" is open-ended, and open-ended actions in a short generation come back rushed or looped.

    Resist the reflex to append "no foot sliding, no skating." Negative phrasing is a weak lever — on models that run without classifier-free guidance the negative field may do nothing at all, and even where it functions it tends to leak the negated concept back in. Describe planted feet rather than prohibiting sliding ones.

    Naming the surface does double duty.

    A named surface is the highest-value single word in a walking prompt, and it pays twice.

    Visually, a surface implies a contact behaviour. Gravel scatters and compresses. Wet asphalt reflects and holds a footprint for a moment. Floorboards are hard and flat and give nothing. Sand deforms under weight. Each gives the model a concrete relationship between shoe and ground instead of an abstract "walking near the floor," and the concrete version drifts less.

    Audibly, it gives you the foley. Models with native audio generate ambience and effects in the same pass as the frames, and a named surface tells that pass what a footstep should sound like. Name two foley events at most, specific and countable — "boot heels on wet asphalt", "a car passing behind" — and let the model fill the ambient bed.

    Get the surface right and audio and visual reinforce each other; leave it out and you have a silent-feeling walk on an abstract floor.

    Where walking actually survives

    Prompting improves the odds. Shot design decides them. Shot size is the highest-yield token in a video prompt, and for walking specifically it is close to determinative.

    Shot Risk Why
    Wide, full body, side-on, feet visible Highest Every contact point is on screen, the ground plane is a visible reference, and stride length is directly comparable to translation speed
    Wide, tracking from behind Moderate Feet are visible but the camera moves with the subject, so absolute ground speed is much harder for a viewer to judge
    Medium, waist up Low No contact points in frame at all; the walk is carried by torso bob and arm swing, which models handle well
    Close-up, walking toward camera Low Minimal ground reference, and the dominant motion is scale change rather than translation
    Insert, feet only Moderate Contact is the subject, so it must be right — but the shot is short, the framing is tight, and one clean step is a much smaller ask than a full cycle

    The pattern is straightforward: risk scales with how much ground-plane reference is visible. The side-on full-body walking shot is the classic cinematography choice and the worst one here, because it puts the exact failure at frame centre with a reference line running through it.

    On length, the general three-to-five-second ceiling applies with extra force. A walk cycle is periodic, so an error in the cycle repeats visibly rather than passing once. Two or three seconds — a handful of steps at an ordinary pace — usually reads better than eight seconds of the same walk, and it matches the cut rate short-form editing wants anyway. A character who has to cross real distance needs several shots and a cut, not one long generation. The wider version of that argument is in where motion coherence breaks as clips get longer.

    Designing around it

    The cheapest fix is not generating the problem.

    Cut on the step. A cut landing on a footfall hides the join and lets you skip the parts of the cycle you did not get. Two two-second clips cut on a heel strike read as continuous walking far more reliably than one four-second clip.

    Start and end mid-stride. Accelerating from standing and decelerating to a stop involve weight shifts much harder than steady-state gait. Enter with the walk established and leave before it resolves.

    Let them arrive. Most shots that "need walking" need the character to be somewhere else by the end. Stepping into frame and stopping gives you the narrative movement with one or two contact events instead of a cycle.

    Use motion blur and depth of field. Shallow focus that leaves the feet soft removes the sharp detail where the error is visible — the same move that works on hands and other high-risk surfaces, covered in hands, teeth and signage.

    Fix the camera relationship. A locked-off camera with a subject walking through frame makes ground speed measurable; a camera tracking alongside at roughly the subject's speed does not. Parameterised camera control is worth reaching for, since holding relative speed constant across a sequence is something prose cannot do repeatably.

    Checking it before it ships

    Watch the feet at full speed, then at reduced speed, then frame by frame through one complete stride. Three passes because the failure hides differently at each rate: at speed you catch gross sliding, at reduced speed you catch the gait resetting, and frame by frame you catch a foot that translates a few pixels while it is supposed to be planted.

    At 25 fps — the default frame rate for video generations here — a single step is roughly a dozen frames, so stepping through one is a matter of seconds. The editor's pre-publish check covers the wider list, and generation duration is the setting to reach for when the answer is "make it shorter."

    FAQ

    Which models handle walking best?

    Ask that about your own footage rather than taking a leaderboard's word for it, because gait quality varies with shot size and subject in ways aggregate scores do not capture. What generalises is the technique. If you are choosing a model for short, high-motion clips, the short-clip shortlist is a starting point, and the agent chat can fan the same walking prompt across several named models in one request so you compare on identical text.

    Does raising the motion setting help?

    Usually not, and often the reverse. Motion level controls how much movement the model introduces, not how physically correct it is — turning it up on a walk cycle gives you a more energetic version of the same sliding. Contact is a specification problem, not an amplitude problem.

    Why does the gait look right for two steps and then break?

    Because periodic motion has to survive being re-derived every frame, and small errors accumulate until the cycle no longer closes. Two or three clean steps is a realistic target for a short clip; a long unbroken cycle is not, and the answer is a cut rather than a longer generation.

    Can I fix sliding in post?

    Not really. Sliding is a mismatch between two motions the model rendered together, so there is no separable layer to correct — retiming changes both at once. Masked regeneration of the lower legs is possible and rarely worth it. Re-shoot shorter, tighter, with contact specified.