Cloth and hair motion that holds together
Hair and fabric are where temporal coherence breaks first. Prompt cues for weight and stiffness, shot choices that hide it, and when to mask and regenerate.
Play a generated clip at quarter speed and watch a strand of loose hair against a bright background. It shimmers — strands appearing, dissolving and reappearing somewhere adjacent, frame after frame. Do the same with a loose sleeve and you see folds that reorganise between frames rather than travelling, a hem that changes length slightly, a seam that migrates.
Nobody watching at speed articulates this. They register the shot as slightly synthetic and cannot say why. Hair and fabric are almost always the reason, because they are the first surfaces where temporal coherence gives out.
Why these two break first
Everything else in a shot has structure holding it steady. A face has bone underneath. A chair has fixed geometry. An arm has a skeleton and a limited range of motion. When the model re-derives those frame to frame, the space of plausible answers is narrow.
Hair and cloth have no skeleton. A head of hair is thousands of thin, near-identical elements whose individual positions are not determined by anything visible — swap two strands and the image is equally valid. Fabric folds in an effectively unbounded number of ways, and the fold pattern in frame twenty does not determine the pattern in frame twenty-one. Both are high-frequency detail with weak frame-to-frame constraints, which is exactly where a model produces something plausible per frame and inconsistent across frames.
That is why the failure looks like shimmer rather than like a mistake. Each individual frame is fine. The sequence is what breaks, which makes it a temporal consistency problem rather than a quality problem, and no amount of resolution fixes it.
Prompt for weight and stiffness, not for beauty
The instinct is to describe hair and cloth aesthetically — flowing, silky, billowing, windswept. Those words specify appearance and say nothing about physical behaviour, and behaviour is what has to stay consistent across frames.
Replace them with weight and stiffness.
| Instead of | Write |
|---|---|
| "long flowing hair" | "heavy hair, moving as a single mass, tips lagging behind the head turn" |
| "silky hair" | "hair falling straight and settling quickly, minimal bounce" |
| "windswept hair" | "damp hair, clumped in strands, barely moving" |
| "a billowing coat" | "stiff waxed canvas coat, holding its folds, swinging from the shoulders" |
| "a flowing dress" | "light cotton, moving a beat after the body, hem settling between steps" |
| "fabric blowing" | "heavy wool that swings rather than ripples" |
Two properties do the work. Weight determines lag: heavy material moves after the body and settles quickly, light material moves with it and keeps moving. Stiffness determines fold count: stiff material holds a few large folds, soft material produces many small ones. A model given "heavy, stiff, holds its folds" has a much narrower set of correct answers per frame than one given "flowing," and a narrower set is a more consistent one.
There is a second-order benefit: fewer, larger folds means fewer high-frequency elements to keep coherent.
Name the force.
Cloth and hair do not move on their own; something moves them. If the prompt does not say what, the model invents a force and it can invent a different one at different points in the clip — which is a large part of why fabric sometimes appears to be reacting to wind from two directions at once.
State the source, the direction, and roughly the strength:
- "steady wind from camera left, strong enough to lift the hem but not the hair"
- "no wind; the coat moves only from her own stride"
- "still air, interior; hair moves only when she turns her head"
That last form is the most under-used one. Interior shots have no wind and models frequently add some anyway, because the training distribution is full of dramatic outdoor footage. Saying that the air is still is the cloth equivalent of stating that the camera is motionless — omitting movement language does not reliably get you stillness, you have to ask for it.
Shot and lighting choices that hide the problem
Prompting reduces the failure. Shot design decides whether anyone can see it.
Tie the hair back. Immediately removes most of the risk. A ponytail is one mass with a defined pivot; loose hair is thousands of independent elements. If the character's look permits it, this is the single highest-leverage decision available.
Prefer structured garments. A jacket, a collared shirt, a coat with a defined shoulder line. Loose scarves, long unstructured dresses and layered chiffon are the hardest cases in the wardrobe.
Watch the backlight. Hair shimmer is most visible when individual strands are separated against a bright background — the rim-lit halo that looks beautiful in a still is the exact condition that exposes strand-level inconsistency in motion. If you want a rim light on a character with loose hair, keep the shot short. Lighting language works best when it names source size, direction and colour temperature rather than mood, and direction is what decides whether the risky detail is separated or absorbed.
Use depth of field. A soft background behind moving hair removes the sharp detail where the eye catches the flicker — the same tactic that works on hands and other high-risk surfaces, covered in hands, teeth and signage.
Keep clips short. Three to five seconds is the working range for temporal stability, and cloth benefits from the low end of it.
Avoid the whip. Fast head turns and sudden direction changes are where hair does its worst work, because the model has to invent a large displacement between frames. A slower turn is easier to render coherently.
Frame rate and the finishing order
Two mechanical points that fix more shimmer than any prompt change.
Do not generate at high frame rates hoping for smoothness. Generated footage sits most comfortably in the 24-to-30 fps band, and video generations here default to 25 fps, inside it. Asking a model for many more frames over the same duration gives it more frames to keep in agreement, which is the thing already failing — so it often reads as more flicker, not less. Interpolate up afterwards if a delivery needs it. Frame rate is a coherence dial, not a quality dial.
Deflicker before you upscale, never after. The order is clean, then 1080p, then step to 4K. Going straight to 4K enlarges every shimmering strand faithfully, and upscaling has no way to know the detail it is amplifying was inconsistent to begin with. Cleanup tools come before upscaling to 4K, not after.
Colour and exposure also drift between clips, and matching them at assembly is standard rather than optional. Shimmer, though, is firmly on the wrong side of what grading can fix — the boundaries are in what a grade can and cannot rescue. A grade adjusts values; it does not make frame twenty-one agree with frame twenty.
When to mask and regenerate
Some shots are worth saving. The tool is regional regeneration: mask the failing area and re-generate only that region rather than re-rolling the whole clip.
This is a reasonable move when:
- The failure is confined to one clearly bounded area — a sleeve, a scarf, the back of the head.
- The rest of the shot is genuinely good, including the parts that are expensive to reproduce, like a specific facial performance.
- The masked region does not overlap the subject's face or hands, where a regional pass risks introducing an identity shift worse than the original problem.
It is the wrong move when the failure is spread across the frame, when re-shooting the clip costs little, or when the region moves so much that the mask itself becomes the hard part. Video segmentation makes a moving mask tractable and inpainting is the repair operation itself. One honest caveat: regional regeneration on video is fiddlier than the still equivalent. Reach for the mask when the shot contains something you cannot easily get again; re-shoot when it does not.
Teams running open models in their own pipelines get one more dial, because self-hosted sampler settings expose the trade between how much motion a model introduces and how much detail it resolves per frame. That is the same trade this whole post is about: calmer shots really are more coherent, and the coherence is bought with movement.
FAQ
Does a character reference stop hair from flickering?
No. A reference binds appearance — who this is, what their hair looks like — and shimmer is a frame-to-frame consistency failure, which is a different axis. A reference will keep the hair the right colour and length while it flickers. Use references for identity and the levers in this post for coherence.
Is this worse on some models than others?
Yes, and it moves with releases often enough that a ranking written today ages badly. Test it directly: generate the same short clip with loose hair and a loose garment on two or three candidates, then step through frame by frame. It is a fast test and it discriminates between models much better than a general quality impression does.
Why does the fabric look fine in the first second and wrong later?
Because coherence degrades as the generation runs and errors accumulate. The first second is closest to the conditioning input and therefore the most constrained. That is the same reason short clips hold better than long ones, and it is a strong argument for trimming a clip to its good opening rather than trying to rescue the tail.
Should I turn the motion setting down?
Worth trying, as long as you know what it does. Motion level controls how much movement the model introduces, and less movement means less to keep coherent — so lowering it often reduces shimmer, at the cost of a shot that may read as static. The better version of the same trade is to keep motion where you want it and reduce what is in the frame that has to move: tied-back hair, structured clothing, still air.