The mouth goes soft after a lip-sync pass
Older lip-sync models rebuild a tiny mouth crop and paste it back soft. Finish with a face restore, or switch models when the patch is structural.
Pause on a frame. The eyes are sharp, the pores are there, the shirt holds a weave. The mouth is a different photograph: a slightly milky oval, teeth a hint too smooth, a faint rectangle you only see because you are looking for it. That is not a grade. That is a crop.
Older lipsync stacks (the Wav2Lip lineage is the one operators still meet) do not repaint the face at the source resolution. They cut a small mouth patch, often 96×96, generate the viseme at that size, and paste it back onto a 1080p or 4K frame. The rest of the picture never went through that bottleneck. The mouth did. You are looking at an upscaled stamp.
A soft island in a sharp frame is a crop, not a grade
Three checks that separate this from "the whole clip is a bit soft":
- Zoom to 100% on a still frame, not on playback. Playback hides the patch. A still makes the resolution step obvious.
- Compare mouth pixels to cheek pixels. Same lighting, same compression. If the cheek holds detail the lip does not, the generator never had those lip pixels.
- Look for the blend edge. A slight halo or a change in grain around the jaw is the paste. Heavy temporal smoothing tries to hide it and instead makes the whole lower face look airbrushed.
This is distinct from accumulating lip-sync drift (a clock / export problem) and from a 1–3 frame lag on plosives (a timing problem). Softness can sit on a perfectly timed mouth. Timing the stamp does not sharpen the stamp.
It is also distinct from a low-resolution source. If the whole face is soft, upscale the clip or reshoot. If only the mouth is soft, the source was fine and the lipsync pass threw resolution away.
Talking-photo models can produce a related look for a different reason: they invent the whole head from a still, so there is no high-resolution original mouth to paste onto. The failure is then "the whole face is a generated texture," not "a postage stamp on a sharp plate." Treat those as a model-choice problem, not a restore problem. Fabric-class photo-to-talking pipelines live in that second bucket.
Why a 96×96 mouth cannot survive a 1080p paste
A 1080p talking-head close-up gives the mouth region on the order of a few hundred pixels across. A 96×96 generator has to invent the rest on the way back. There is no information in that crop for individual teeth edges, lip texture, or stubble that crosses the vermilion border. The paste is a low-pass filter with a viseme on it.
What you will see, in order of how often it bothers a QC pass:
| Artifact | Cause | What does not fix it |
|---|---|---|
| Milky lips, plastic teeth | Patch resolution | Color grade, contrast, a different caption style |
| Soft rectangle around the mouth | Blend of the upscaled crop | Sharpening the whole frame |
| Jaw / neck seam | Crop did not include enough jaw | Another pass of the same model |
| Mouth sharp on some frames, mushy on others | Tracking jitter on the crop | Higher bitrate export |
Sharpening the whole frame is the trap. It will crisp the cheeks and the background and amplify whatever the paste already did: grain mismatch, flicker on the teeth, a boiling outline. Over-sharpening in an upscaler is how a mild stamp becomes a sparkling one. Soft is better than sparkling. Do not "fix" a resolution step with a detail slider.
A face-restore finish for mild cases
When the timing is good and the patch is only a little soft (visible at 100%, not at feed distance), a restore pass is cheaper than a new lipsync.
Still or talking-photo output. Run a face-aware image upscale on a representative frame first. If the mouth holds at 2× without inventing a second row of teeth, run the same treatment as your finish. Two 2× passes beat one 4× pass. Check text and teeth at full size: those are where generative upscalers fabricate.
Video re-sync output. Use a video upscale, not an image upscaler looped over every frame. Frame-by-frame image restore is how you get a mouth that sparkles: each frame invents a slightly different tooth edge. Video upscalers exist because the invented detail has to agree with its neighbours.
Rules that keep the restore from making things worse:
- Restore after you have accepted the sync. Do not restore, then re-time, then restore again.
- Do not feed a heavily compressed download. The upscaler will treat blocking as pores and enlarge it.
- Stop if teeth start to shimmer, if a second lip line appears, or if the blend rectangle becomes easier to see. Those are signs the model is drawing structure that was never there.
- Judge at the delivery size. A patch that is obvious at 200% on a 27-inch monitor is often invisible on a phone feed. Spend the restore on close-ups that will be inspected, not on every medium shot.
The restore is a polish on a stamp that is almost good enough. It cannot put 1080p of real mouth back into a 96×96 generate. If the rectangle is obvious at 1×, skip this section.
When to switch models instead of polishing
Switch when any of these is true:
- The soft patch is visible without zooming.
- The blend edge crawls along the jaw when the head turns.
- Teeth flicker, even slightly, on a smile.
- You already tried a restore and the patch is still a different resolution from the cheeks.
Those are structural. The crop was too small, or the model never stopped being a crop-and-paste model. A stronger re-sync model regenerates the lower face at something closer to source resolution, with more of the original pixels as context.
On Versely, that usually means moving the job to a current video-to-lipsync model rather than running the same cheap pass twice:
- Sync Lipsync 2.0 for close-ups, paid placements, and anything where the mouth fills the frame. It is a video-to-lipsync model: existing footage plus new audio, with sync-mode control (
cut_off,loop,bounce,silence,remap). - Kling Lipsync when you are re-voicing or translating a talking-head clip and need a dedicated video-to-lipsync path rather than a talking photo.
- VEED Lipsync when you want the fast, budget-friendly video-to-lipsync path and the shot is a medium. Do not spend Sync 2.0 credits to polish a medium shot that a restore would have handled.
Photo-input jobs stay on generate_lipsync with a still plus audio. If your source is already footage, pick a video-to-lipsync model rather than a talking-photo one. Match the input shape to the job.
A request that avoids the cheap-stamp-then-sharpen loop:
"Re-sync this talking-head clip to the WAV, not the MP3. Use Sync Lipsync 2.0 because this is a close-up. If the mouth comes back as a soft patch against a sharp face, do not sharpen the whole frame. Tell me, and we'll either upscale the clip as video or accept that this source angle is past what a crop can hide."
Source still matters. Profile, a hand on the jaw, a dense moustache: every model, old or new, has less mouth to work with. Switching models does not invent the far side of a 90° face. It only stops you paying twice for a 96×96 paste.
FAQ
Can I just export at a higher bitrate to sharpen the mouth?
No. Bitrate preserves what is in the file. The stamp was already low-resolution when it was pasted. A fatter encode of a 96×96 viseme is a fatter encode of a 96×96 viseme.
Will upscaling the original clip before lipsync help?
Sometimes, on a genuinely low-res source. It will not help if the source is already 1080p and the softness appears after the lipsync pass. In that case the generator discarded the resolution. Upscale the result (as video), or switch models.
Why do some frames look fine and the next look mushy?
The tracker moved. On a given frame the crop caught the mouth cleanly; on the next it included more cheek or clipped a lip, and the 96×96 generate had to fill. That jitter is another reason not to stay on a crop-and-paste model for close-ups. Temporal smoothing will hide it by blurring more frames, which is how "a few mushy frames" becomes "a soft lower face."
Is a talking photo supposed to have this rectangle?
Not the rectangle. A talking photo has no original mouth pixels, so the whole head is generated texture. If you see a local milky oval on otherwise photographic footage, you are looking at a video re-sync stamp, not a talking photo. Use a current re-sync model. Keep Fabric-class models for the still-image job they were built for.