Lipsync when the mouth is the product
Talking-avatar examples in the app. Use them when the deliverable is a mouth saying the line, not a world.
A talking-avatar example is a mouth saying a line. That is the whole deliverable. The room, the weather, the camera move — none of that is the job. If the brief is a world, you are in the wrong studio.
AI lipsync is the surface: a face plus a script or an audio file, out as a talking head. Avatar X when the presenter is the file is the routing rule for a stock presenter you type a script at. The Fabric / Hedra / sync.so split is in VEED Fabric vs Hedra vs Sync. This post is the job test that sits in front of all three: is the mouth the product, or is the world?
The job test
Hold the brief up to one sentence.
"We need someone to say this, on camera, and the someone is a file."
If that is the brief, lipsync is the tool. Photo or avatar in, line in, talking clip out. UGC ads, course explainers, personalised outreach, a digital twin reading a weekly update — the product is the mouth. A generate-a-world model will give you a cinematic kitchen and a face that cannot be trusted to finish the sentence.
"We need a place, a product, a motion, a feeling."
If that is the brief, do not open lipsync. Send it to image-to-video or text-to-video. A talking-head engine will not invent a decent turntable, and you will waste the take trying to prompt weather into a portrait model.
The failure mode is mixing them on one shot: generate a world, then pretend the mouth in that world is a presenter you can re-line. It is not. Re-lining a generated extra is a different, worse job than starting from a still of a face you control.
What "the mouth is the product" actually means
It means the viewer is here for the words, delivered by a face they can look at. Three honest versions:
| Deliverable | Input | Why lipsync |
|---|---|---|
| Photo-born UGC | One still + a script | No shoot. The face is the ad. |
| Stock avatar presenter | A catalogue avatar + a script | The presenter is a file. Avatar X belongs here. |
| Re-sync of an existing take | Video + new audio | The picture is already approved; only the mouth has to move. |
It does not mean "any clip with a person in it." A street interview, a product demo where hands do the talking, a brand film where the founder walks a floor — those have mouths in frame. The product is the scene. Driving those through a talking-head engine flattens them into a teleprompter with a background.
Two more tells that you are in the mouth job:
- You would ship the clip if the background were a wall.
- You would not ship the clip if the mouth missed a plosive.
If both are true, pay for lipsync. If only the first is true, you have a talking still you could have captioned. If only the second is true, you have a film with a dialogue problem, which is a Veo-class audio job, not a lipsync job.
Pick the engine from the input, not the vibe
Once the job is "mouth," the input decides the engine. Do not pick the one whose example looked nicest in a vacuum.
One photo and a script, no footage. Fabric: still + text, voice and motion from one generate. Front-facing, even light, face large in frame. Sunglasses, heavy shadow, or a small face in a wide shot will cap the output. The Fabric comparison is the bake-off; the input is what actually moves quality.
A catalogue presenter, no specific human. Avatar X. Script in, stock avatar performs it. Do not send a product still or a "cinematic world with a spokesperson" brief here.
Existing video, new language or a replacement read. Re-sync / dub. If the deliverable is the same person in French, you are in the dubbing chain, not in a new avatar.
Stylised character, budget is the constraint. Hedra: cheaper, slower, with known failure modes (over-smiling, head tilt, missed plosives). Not a Fabric substitute.
The app's talking-avatar examples — HeyGen-class twins, Fabric still-to-talker — exist so you can watch a mouth say a line before you spend. If you find yourself grading the background, you have already left the job.
What you still have to do after the mouth works
Lipsync is not a finished post. It is a talking plate.
- Consent. A cloned or animated real person needs a written grant. A generated face does not get you out of an impersonation problem if the viewer would think it is someone specific.
- Captions. A muted feed still has to read the line. Burn them in; do not trust the platform to style a talking-head ad.
- The rest of the ad. Proof, offer, end card. A mouth saying a claim with no product in frame is a talking thumbnail. Cut the product in, or do not run it as a product ad.
If the brief grows a second shot — pack, overlay, location — generate that shot on a video model and cut. Do not ask the lipsync engine to become a world.
FAQ
Can I use lipsync for a product demo if the person is holding the product?
Only if the face is still the shot and the product is a prop in frame. If the demo is hands, pour, texture, or a turntable, the mouth is not the product — the object is. Generate or shoot the demo. Add a talking bumper before or after if you need a line.
Is a talking avatar a substitute for a founder talking-head?
It is a substitute for a presenter you do not have. It is not a substitute for a founder the audience already knows. Viewers who follow a person will notice a twin. Viewers who have never met you will accept a catalogue face if the line is good and the disclosure is honest. Do not run the twin in the founder's slot and hope.
Why not generate the whole talking scene with a dialogue video model?
Because that model is building a world and hoping the mouth keeps up. Always-on dialogue is a different job, with a different failure mode (the room drifts, the identity drifts, the line gets paraphrased). When the deliverable is the line, start from a locked face and move the mouth. When the deliverable is the scene, do not start in lipsync.
Do I need a real photo, or will a generated face work?
Either, with different rights. A real photo needs consent from the person in it. A generated face needs to not be a living person you did not clear. Front-facing, even light, head-and-shoulders either way. The engine will not rescue a bad crop, generated or shot.