Lip and SFX Sync Tolerance in Finishing
Hold lips to ±1 frame on film and ±2 on broadcast. Slip generated SFX inside that window, and regenerate only when a slip would break the rest of the scene.
A mouth that is one frame late is a finishing problem. A whoosh that is two frames late is usually a slip. Regenerating because something "feels off" is how you spend another generation on a 40-millisecond nudge.
This is not a lipsync model-failure note. When a mouth starts right and walks away over twenty seconds, that is an export clock, covered in lip sync that drifts across the clip. Finishing is the other case: picture and sound are locked to the same rate, the offset is constant, and you are deciding whether to slip a clip, slip a whole track, or pay for a new render.
Frames are the house unit. Milliseconds are the standard.
Finishing rooms still talk in frames. A common house rule is ±1 frame on film (24 fps) and ±2 frames on broadcast (25 fps, or 29.97 where one frame is 33.4 ms). Those are not the perception numbers, and they are not a licence to park every hit two frames late. Convert them before you treat them as law.
| Rate | 1 frame | House ±1 frame | House ±2 frames |
|---|---|---|---|
| 24 fps | 41.7 ms | ±41.7 ms | ±83.3 ms |
| 25 fps | 40.0 ms | ±40.0 ms | ±80.0 ms |
| 30 fps | 33.3 ms | ±33.3 ms | ±66.7 ms |
Versely's default video frame rate is 25 fps, so one frame on a Versely timeline is 40 ms. Mixing 24 fps clips onto that timeline without a stated conform is its own error; 24 fps clips on a 25 fps timeline is the arithmetic.
Named standards sit next to the house rule, not underneath it.
ITU-R BT.1359-1 is a perception document. Detectability is about +45 ms (sound ahead of picture) to −125 ms (sound behind); acceptability is about +90 ms to −185 ms. A positive value means sound is early. Viewers notice early sound sooner than late sound.
EBU R37 is a production budget: +5 ms / −15 ms at any one stage, and +40 ms / −60 ms at an output intended for emission. At 25 fps that overall budget is +1.0 / −1.5 frames, tighter than a casual "±2 frames" on the lead side. A film-exchange figure of ± half a frame is about ±22 ms at 24 fps. A one-frame theatrical slip is already past that note.
So the house rule is a starting grid, not a QC stamp. Use it this way:
- Talking heads and on-camera speech: hold inside ±1 frame, and prefer sound late over sound early if you must miss zero. +80 ms (two frames early at 25 fps) is already past BT.1359's detectability edge.
- SFX, hits, footsteps, whooshes: ±1 frame is the working target. ±2 frames is the last stop before you admit the file is wrong.
- Do not use BT.1359's −185 ms acceptability figure as a finishing tolerance. That number is "the viewer will still watch the programme," not "this door slam is in sync."
How to measure the offset before you touch a fader
Guessing by ear in a laptop speaker is how a one-frame error becomes a three-frame "fix." Scrub, then count.
- Pick an event you can see and hear: a plosive (
p,b,t), a door close, a clap, a foot plant, a glass put down. - Park on the picture frame where the event is unmistakeable (lips close, latch meets strike, heel hits).
- Park on the audio peak or transient for the same event.
- Count frames between them, and write the sign: sound early (+) or sound late (−).
- Repeat at the start and at the end of the clip.
Two readings, same sign, same size: constant offset. Slip. Two readings, same sign, growing size: drift. Stop slipping and re-export; a slip cannot fix a clock. Two readings, different signs: you are measuring two different events, or the performance itself is late. Do not average them into a "±1.5 frames" and call it finished.
A plosive is a better lip event than a vowel. A latch click is a better SFX event than a rumble. If you cannot find an event that exists in both picture and sound, you do not have a sync measurement. You have a taste note.
Slip the SFX. Do not slip the scene.
Generated sound effects arrive as files you can nudge. Generate a sound effect from a prompt, drop it on the timeline, and treat it as a clip with a start. The cheap move when a hit is one or two frames off is to slip that clip. The expensive move is to regenerate the video so the door closes on a different frame, or to regenerate the effect in the hope the model lands on the picture by accident.
Slip when all of these are true:
- The offset is constant.
- The event is discrete (a hit, a whoosh, a footstep, a riser that must land on a cut).
- Sliding the audio does not pull a different on-screen event out of sync. A door slam that also contains the room's tone under the rest of the shot is no longer discrete.
- The new position is inside the house window, and not early on a talking face.
Regenerate (or run a dedicated lipsync pass) when any of these are true:
- Speech is off by more than one frame and you cannot slip the voice without unsitting the rest of the production audio.
- The mouth is wrong in shape, not just in time. A late mouth that is also the wrong viseme will not be saved by a slip.
- The SFX is baked into a native-audio render and there is no separate file to slip. Extracting or replacing the bed is then a mix decision, not a nudge.
- The offset grows through the clip. That is not a tolerance problem.
A constant one-to-three-frame lag on a talking head can be model lag or encoder delay. A slip of the voice is legal only if nothing else on that track has to stay glued to picture. If the same track holds footsteps that were right, split it: slip the voice, leave the feet. If you cannot split it, you do not have a slip. You have a rebuild.
When the mouth actually needs a new pass, the lipsync video task is the generation surface. Finishing still happens after that pass, on a timeline that can slip by frames. Credits go to the generation. The slip costs nothing in the NLE.
A whoosh that leads a cut by a frame or two is often intentional. J-cuts and L-cuts with generated audio are editorial, not errors. Do not "correct" a J-cut back to zero because a sync overlay turned red.
A slip/regen cheat sheet
| What you see | First move | Regen if |
|---|---|---|
| Door slam 1–2 frames late, SFX is its own file | Slip the SFX earlier | The transient is the wrong object (a wood knock for a metal latch) |
| Whoosh 1 frame early on a cut | Leave it, or slip 1 frame later | It masks a word |
| Footsteps wander around the plants | Slip per step, or replace the file | The gait in picture is itself broken |
| Mouth 1 frame late, offset constant, voice is a separate track | Slip voice later by 1 frame (prefer late to early) | Other glued SFX live on the same track and you cannot split |
| Mouth 2+ frames late on a close-up | Lipsync pass or new take | A 2-frame slip on a close-up still reads as late to you on a large monitor |
| Mouth right at 0:01, late at 0:12 | Re-export at native rate | You already re-exported and it still walks |
Prefer sound slightly late over slightly early when you are inside a frame and have to choose. Early sound on a face is the one viewers name.
Do the QC on speakers that can show a transient, at the delivery frame rate. Versely's editor can assemble and preview at 480p with a short per-user cooldown; the final export is charged once. That preview is for pace, not sync. Count frames on the finishing timeline.
FAQ
Is ±2 frames good enough for a talking head?
As a house ceiling on broadcast, it is a last stop, not a target. At 25 fps, two frames early is 80 ms, past BT.1359's +45 ms detectability. Two frames late is 80 ms, inside the −125 ms detectability figure and outside EBU R37's −60 ms emission budget. For a close-up, finish at ±1 frame and prefer late to early.
Should I slip picture or sound?
Slip sound. Picture is what other departments already signed. Moving picture to chase a whoosh will also move lips and titles. The exception is a clip whose baked action is late against a locked music hit: slip the clip as a whole, picture and its native sound together.
Does Versely's 25 fps default change the tolerance?
It changes the conversion. One frame is 40 ms, which happens to equal R37's +40 ms "sound before picture" overall limit. It does not make ±2 frames safer. If you deliver 24, recount. Do not import a 25 fps habit onto a 24 fps master without doing the millisecond column.
When is a regen actually cheaper than a slip?
When the event you would slip is not separable: speech baked into a native-audio take, an SFX that is the room for the whole shot, or a mouth that is the wrong shape. A slip of a discrete file costs a drag. A regen costs another generation. Use the regen when the slip would desync something you already approved.