Lip sync drifts across the clip
Most lip-sync drift is an export fault, not a model fault. Re-export at native frame rate, skip optimized rendering, and feed uncompressed audio.
If the mouth is right on the first line and a half-beat late by the last line, regenerate is the expensive guess. The cheap one is the export. Frame-rate assumes, container timestamp rounding, an NLE writing from optimized media, and MP3 encoder delay will all walk a mouth off its audio without the lipsync model having done anything wrong.
Practitioners who actually ship talking heads report the same cluster: re-export at the clip's native frame rate, turn off optimized rendering, and feed uncompressed WAV rather than MP3. That trio resolves the large majority of severe drift cases without another generation. The remaining cases are real model lag, and they look different. They start a frame or two late and stay there. They do not accumulate.
If the mouth starts right and then walks away, look at the export
Two shapes, two jobs.
Accumulating drift. Sync at 0:02, visibly off at 0:12, comic at 0:25. The error grew with time. That is a clock disagreement: the picture is playing at a different rate than the audio, or the file's timestamps are being rounded every frame in a way that does not cancel. Models do not usually do this. Exporters do.
Constant offset. The mouth is one to three frames late from the first plosive to the last. The gap does not grow. That can be the model (especially older crop-and-paste stacks on fast speech), or it can be a fixed audio delay from a compressed encode. A constant offset is a nudge. Accumulating drift is a re-export.
Before you spend another credit on Sync Lipsync 2.0 or Kling Lipsync, scrub two plosives: one in the first second, one in the last three. If the first lands and the second does not, stop regenerating.
Versely's default video frame rate is 25 fps, not 24. A lot of generated clips, phone takes, and "30 fps" exports are not 25. One extra or missing frame per second is 40ms of picture/audio disagreement per second. That is already a mouth problem on a 10-second clip, and it is the same arithmetic as 24 fps clips dropped on a 25 fps timeline. Lip-sync makes it impossible to ignore.
The three export settings that cause most of the drift
1. Native frame rate, not the sequence rate
The clip has a rate. The sequence has a rate. The export dialog has a rate. If those three disagree, one of them is being conformed, and conform without an explicit policy retimes picture against audio.
Export at the rate the lipsync file was written. If the model or the camera delivered 25 fps, do not "help" it to 30 on the way out of Premiere, Resolve, or a web exporter. A 25-to-30 conform is a 1.2× stretch of the frame grid. Audio that was aligned to 25 fps picture is now playing against a different set of frames.
If the project truly has to deliver 30, conform the lipsync clip first, with a stated policy (duplicate, interpolate, or retime picture and audio together), then check a plosive at both ends, then export. Silent conforms in the export dialog are how drift ships.
2. Optimized rendering off
NLEs keep a fast playback copy: optimized media, proxies, render files, "use previews." Those copies are allowed to have a different rate, a different GOP structure, and rounded timestamps. They are for scrubbing. They are not for delivery.
On export, disable the path that writes from those copies. In Premiere-family dialogs that is the "use previews" / optimized-media option. In Resolve it is "use optimized media if available." Whatever the label, the intent is the same: the deliverable should be encoded from the original media, at the original rate, not from a playback cache.
This is the setting people toggle on for a late deadline and never toggle off. It is also the setting that makes a lipsync QC pass in the viewer disagree with the file you uploaded.
3. Uncompressed audio in
MP3 (and AAC) add encoder delay. The codec inserts priming samples so the first audible frame has the right context. Players differ in how they honour that delay. The result is a small, often constant, offset: the mouth starts on time in the editor and a few tens of milliseconds late in the MP4. On a talking head, tens of milliseconds is a viseme.
WAV (PCM) does not do this. Feed the lipsync job WAV. Export with PCM or a well-behaved AAC encode from a clean WAV, not from an MP3 you bounced "to save size." If the voiceover started life as MP3, transcode to WAV once, sync against that, and do not go back.
A compressed voiceover under a generated mouth is a common pipeline because TTS downloads often arrive as MP3. Transcode before the lipsync call, not after.
A checklist that does not regenerate the clip
Run this in order. Stop at the first step that puts the last plosive on the word.
- Confirm the container rate.
ffprobeonr_frame_rateandavg_frame_rate. If they disagree, you have VFR; transcode to constant rate first. Caption and lipsync both assume a grid. - Re-export at native fps. Same number the lipsync file already has. No 23.976 to 25, no 25 to 30, no "match sequence."
- Export from original media. Optimized rendering off, proxies off, "use previews" off.
- Replace the audio with WAV. If the timeline still has the MP3, swap it. Re-export. Do not regenerate the mouth to chase an encoder delay.
- Watch the tail, not the hook. QC a plosive in the last 10% of the clip. Accumulating drift hides in the first five seconds, which is what everyone reviews.
- Only then regenerate. If the offset is constant after a clean export, it is a model or a source problem. See how lipsync models work for the source-side break conditions (profile, occlusion, music under the vocal). A constant 1–3 frame lag on fast speech is a different article; it is not this checklist.
If you assembled the talking head in Versely's editor rather than an NLE, the same idea still holds: do not mix rates in one EDL, and do not judge sync off a compressed download. The editor's 480p preview pass is free, with a short per-user cooldown; a single charge applies to the final export. Preview is enough to see drift (drift is timing, not resolution). Use it to check the tail before you pay for the export.
A request that keeps the agent from wasting a regeneration:
"Do not re-run lipsync. Confirm this clip's frame rate, re-export at that native rate with the original audio replaced by a WAV of the same take, and I'll QC a plosive at the start and at the end."
If you are still in the generate step, lipsync a photo or video to audio and the AI lipsync tool are the right surfaces. They are not the fix for a file that already synced in the viewer and drifted in the MP4.
When it really is the model
Regenerate when:
- The offset is constant after a native-rate, original-media, WAV export.
- The source is a profile, a hand over the mouth, or music baked into the vocal. Those are input problems. Clean the input (frontal crop, isolated vocal), then generate once.
- The mouth smear is local (a soft rectangle, flickering teeth) rather than a growing gap. That is synthesis quality, not a clock. Switching to a stronger re-sync model is then the right spend. Sync 2.0 is the usual precision pick for close-ups.
Do not regenerate when:
- The first line is locked and the last line is not.
- The editor viewer is in sync and the uploaded file is not.
- You just conformed 24 to 25, or 25 to 30, in the export dialog.
- The voiceover is still an MP3.
Those four are export. They will survive every model you throw at them, because the model never sees the clock you introduced on the way out.
FAQ
Why did the preview look synced and the upload did not?
The preview was playing original media at the native rate. The upload was a conform, a proxy encode, or an MP3 bounce. QC the file you will post, on the device you will post from, not the timeline. A 480p editor preview (free, short per-user cooldown) is for catching this before the charged export.
Does exporting at 60 fps "fix" drift?
No. A higher delivery rate does not reconcile two clocks. If the source is 25 fps, 60 fps is another conform. Stay native unless you have a stated interpolation plan, and even then check a plosive at both ends.
I already mixed music under the dialogue. Could that be the drift?
Music under the vocal confuses the alignment the model ran, which usually looks like a constant mushy offset, not a walk. Isolate the voice, sync, then mix the bed back. If the mouth started right and then separated, you still have an export-rate problem on top.
Can I nudge the audio 80ms and call it done?
For a constant offset, yes. For accumulating drift, a nudge that saves 0:08 will miss 0:20. Measure the gap at two timestamps. If they disagree, re-export. If they agree, nudge.