Output quality

    Temporal consistency

    Also called Frame-to-frame stability, Flicker.

    Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.

    It is the failure mode you cannot see in a still. Pause any frame of a bad clip and it may look excellent; play it and the logo shimmers, the pattern crawls, the wall behind the subject breathes. Each frame was individually plausible and collectively unstable.

    The causes are structural. A model that solves frames independently has no reason to make the same arbitrary choice twice, so anything the prompt did not pin down — the exact weave of a fabric, the position of a background object — is re-decided every frame. Models built for video address this by generating the sequence as one object rather than as a stack of pictures.

    Practically, three things make it worse: longer clips, more motion, and finer repeating detail. Stripes, small text, foliage and crowds are where it shows first, which is why a busy frame is a harder brief than a clean one even when the subject is identical.

    In practice

    • Review at full speed before you review frame by frame — the artefact only exists in motion.
    • Reduce motion, shorten the clip, or simplify the background when a shot will not stabilise.
    • Fine repeating patterns are the canary: if they crawl, everything else is borderline too.

    The mistake to avoid

    Trying to fix flicker with a higher resolution. More pixels give the instability more room; they do not resolve it.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.