Guides

    Lip and SFX Sync Tolerance in Finishing

    Hold lips to ±1 frame on film and ±2 on broadcast. Slip generated SFX inside that window, and regenerate only when a slip would break the rest of the scene.

    Versely Team9 min read

    A mouth that is one frame late is a finishing problem. A whoosh that is two frames late is usually a slip. Regenerating because something "feels off" is how you spend another generation on a 40-millisecond nudge.

    This is not a lipsync model-failure note. When a mouth starts right and walks away over twenty seconds, that is an export clock, covered in lip sync that drifts across the clip. Finishing is the other case: picture and sound are locked to the same rate, the offset is constant, and you are deciding whether to slip a clip, slip a whole track, or pay for a new render.

    Frames are the house unit. Milliseconds are the standard.

    Finishing rooms still talk in frames. A common house rule is ±1 frame on film (24 fps) and ±2 frames on broadcast (25 fps, or 29.97 where one frame is 33.4 ms). Those are not the perception numbers, and they are not a licence to park every hit two frames late. Convert them before you treat them as law.

    Rate 1 frame House ±1 frame House ±2 frames
    24 fps 41.7 ms ±41.7 ms ±83.3 ms
    25 fps 40.0 ms ±40.0 ms ±80.0 ms
    30 fps 33.3 ms ±33.3 ms ±66.7 ms

    Versely's default video frame rate is 25 fps, so one frame on a Versely timeline is 40 ms. Mixing 24 fps clips onto that timeline without a stated conform is its own error; 24 fps clips on a 25 fps timeline is the arithmetic.

    Named standards sit next to the house rule, not underneath it.

    ITU-R BT.1359-1 is a perception document. Detectability is about +45 ms (sound ahead of picture) to −125 ms (sound behind); acceptability is about +90 ms to −185 ms. A positive value means sound is early. Viewers notice early sound sooner than late sound.

    EBU R37 is a production budget: +5 ms / −15 ms at any one stage, and +40 ms / −60 ms at an output intended for emission. At 25 fps that overall budget is +1.0 / −1.5 frames, tighter than a casual "±2 frames" on the lead side. A film-exchange figure of ± half a frame is about ±22 ms at 24 fps. A one-frame theatrical slip is already past that note.

    So the house rule is a starting grid, not a QC stamp. Use it this way:

    • Talking heads and on-camera speech: hold inside ±1 frame, and prefer sound late over sound early if you must miss zero. +80 ms (two frames early at 25 fps) is already past BT.1359's detectability edge.
    • SFX, hits, footsteps, whooshes: ±1 frame is the working target. ±2 frames is the last stop before you admit the file is wrong.
    • Do not use BT.1359's −185 ms acceptability figure as a finishing tolerance. That number is "the viewer will still watch the programme," not "this door slam is in sync."

    How to measure the offset before you touch a fader

    Guessing by ear in a laptop speaker is how a one-frame error becomes a three-frame "fix." Scrub, then count.

    1. Pick an event you can see and hear: a plosive (p, b, t), a door close, a clap, a foot plant, a glass put down.
    2. Park on the picture frame where the event is unmistakeable (lips close, latch meets strike, heel hits).
    3. Park on the audio peak or transient for the same event.
    4. Count frames between them, and write the sign: sound early (+) or sound late (−).
    5. Repeat at the start and at the end of the clip.

    Two readings, same sign, same size: constant offset. Slip. Two readings, same sign, growing size: drift. Stop slipping and re-export; a slip cannot fix a clock. Two readings, different signs: you are measuring two different events, or the performance itself is late. Do not average them into a "±1.5 frames" and call it finished.

    A plosive is a better lip event than a vowel. A latch click is a better SFX event than a rumble. If you cannot find an event that exists in both picture and sound, you do not have a sync measurement. You have a taste note.

    Slip the SFX. Do not slip the scene.

    Generated sound effects arrive as files you can nudge. Generate a sound effect from a prompt, drop it on the timeline, and treat it as a clip with a start. The cheap move when a hit is one or two frames off is to slip that clip. The expensive move is to regenerate the video so the door closes on a different frame, or to regenerate the effect in the hope the model lands on the picture by accident.

    Slip when all of these are true:

    • The offset is constant.
    • The event is discrete (a hit, a whoosh, a footstep, a riser that must land on a cut).
    • Sliding the audio does not pull a different on-screen event out of sync. A door slam that also contains the room's tone under the rest of the shot is no longer discrete.
    • The new position is inside the house window, and not early on a talking face.

    Regenerate (or run a dedicated lipsync pass) when any of these are true:

    • Speech is off by more than one frame and you cannot slip the voice without unsitting the rest of the production audio.
    • The mouth is wrong in shape, not just in time. A late mouth that is also the wrong viseme will not be saved by a slip.
    • The SFX is baked into a native-audio render and there is no separate file to slip. Extracting or replacing the bed is then a mix decision, not a nudge.
    • The offset grows through the clip. That is not a tolerance problem.

    A constant one-to-three-frame lag on a talking head can be model lag or encoder delay. A slip of the voice is legal only if nothing else on that track has to stay glued to picture. If the same track holds footsteps that were right, split it: slip the voice, leave the feet. If you cannot split it, you do not have a slip. You have a rebuild.

    When the mouth actually needs a new pass, the lipsync video task is the generation surface. Finishing still happens after that pass, on a timeline that can slip by frames. Credits go to the generation. The slip costs nothing in the NLE.

    A whoosh that leads a cut by a frame or two is often intentional. J-cuts and L-cuts with generated audio are editorial, not errors. Do not "correct" a J-cut back to zero because a sync overlay turned red.

    A slip/regen cheat sheet

    What you see First move Regen if
    Door slam 1–2 frames late, SFX is its own file Slip the SFX earlier The transient is the wrong object (a wood knock for a metal latch)
    Whoosh 1 frame early on a cut Leave it, or slip 1 frame later It masks a word
    Footsteps wander around the plants Slip per step, or replace the file The gait in picture is itself broken
    Mouth 1 frame late, offset constant, voice is a separate track Slip voice later by 1 frame (prefer late to early) Other glued SFX live on the same track and you cannot split
    Mouth 2+ frames late on a close-up Lipsync pass or new take A 2-frame slip on a close-up still reads as late to you on a large monitor
    Mouth right at 0:01, late at 0:12 Re-export at native rate You already re-exported and it still walks

    Prefer sound slightly late over slightly early when you are inside a frame and have to choose. Early sound on a face is the one viewers name.

    Do the QC on speakers that can show a transient, at the delivery frame rate. Versely's editor can assemble and preview at 480p with a short per-user cooldown; the final export is charged once. That preview is for pace, not sync. Count frames on the finishing timeline.

    FAQ

    Is ±2 frames good enough for a talking head?

    As a house ceiling on broadcast, it is a last stop, not a target. At 25 fps, two frames early is 80 ms, past BT.1359's +45 ms detectability. Two frames late is 80 ms, inside the −125 ms detectability figure and outside EBU R37's −60 ms emission budget. For a close-up, finish at ±1 frame and prefer late to early.

    Should I slip picture or sound?

    Slip sound. Picture is what other departments already signed. Moving picture to chase a whoosh will also move lips and titles. The exception is a clip whose baked action is late against a locked music hit: slip the clip as a whole, picture and its native sound together.

    Does Versely's 25 fps default change the tolerance?

    It changes the conversion. One frame is 40 ms, which happens to equal R37's +40 ms "sound before picture" overall limit. It does not make ±2 frames safer. If you deliver 24, recount. Do not import a 25 fps habit onto a 24 fps master without doing the millisecond column.

    When is a regen actually cheaper than a slip?

    When the event you would slip is not separable: speech baked into a native-audio take, an SFX that is the room for the whole shot, or a mouth that is the wrong shape. A slip of a discrete file costs a drag. A regen costs another generation. Use the regen when the slip would desync something you already approved.