Three ways to make a face talk, and when each wins
A native-audio model, a lipsync pass, or a photo-to-avatar render. Compared on realism, script length ceilings, and what each costs when the copy changes.
There are three ways to get a talking person out of a generation pipeline, and every comparison of them I have read scores the same axis: which looks most real. That axis has been converging for a year and is now the least useful thing to compare on. All three routes produce something that survives a scroll.
The axis that actually separates them is what happens on Thursday afternoon when legal changes one number in the script. On one route that is a re-render of an asset you already approved. On another it is a new performance, new framing and a new round of approvals. Same visual quality, wildly different production cost, and it is entirely determined by which route you picked on Monday.
The three routes
Route A: the model says the line. A native-audio video model generates picture and dialogue in the same pass. You write a prompt describing a person and what they say, and speech comes out attached to a face the model invented on the spot.
Veo 3.1 tops out at 8 seconds per generation. FLUX.3 text-to-video goes to 20. Seedance 2.0 and Happy Horse 1.1 sit at 15. As a rough planning number, natural ad-read pace is around two and a half words a second, so an 8-second ceiling is roughly twenty words and a 15-second ceiling is roughly forty. Anything longer is multiple generations, and multiple generations means the same face has to be re-established each time.
The thing to understand about this route is that you do not select a voice. You describe one, and the model produces a performance that fits the face it also just produced. That coupling is the appeal: mouth, expression, breath and audio all come from one pass, so nothing has to be matched afterwards. It is also the liability, because the same coupling means you cannot change one without the other.
Route B: a lipsync pass over footage you already have. You have a video of a person talking. You have new audio. A video-to-lipsync model rewrites the mouth to match.
Sync Lipsync 2.0 is the precision option, with sync modes (cut_off, loop, bounce, silence, remap) that decide what happens when the audio and the footage are different lengths. That parameter is worth more than it sounds: audio is almost never exactly the length of the clip, and the default behaviour on a mismatch is the difference between a clean deliverable and a re-run. Sync React 1 goes further, driving facial expression and head movement from a chosen emotion, with modes that limit changes to the lips, the whole face, or the head.
Script length here is not a model constraint at all. It is a footage constraint. Twelve minutes of source footage takes twelve minutes of audio.
Route C: a photo becomes a talking head. One still image plus either audio or a script, and the model animates it.
VEED Fabric 1.0 Text is the shortest path: photo in, script in, and the voice is generated to match the character in the image. HeyGen Avatar 4 takes a selfie and a chosen voice. Kling Avatar and Wan 2.2 Speech to Video both take an image plus an audio track; the Wan entry adds body movement and camera work on top of the mouth.
This route has the widest quality spread of the three, and it is almost entirely determined by the input photo. Frontal, evenly lit, unobstructed mouth, neutral expression, and the output is good. Three-quarter angle with hair across the jaw and it is not.
The comparison that matters
| A: native-audio model | B: lipsync pass | C: photo to avatar | |
|---|---|---|---|
| You supply | A prompt | Footage plus audio | A photo, plus audio or a script |
| Who chooses the voice | The model | You | You, mostly |
| Script length ceiling | The duration ladder, 8 to 20 seconds | Your footage | Generous, model-dependent |
| Identity across clips | Re-established each generation | Inherited from the footage | Fixed by the photo |
| Best-case realism | Highest, everything from one pass | High, limited by source footage | Good to high, set by the photo |
| Copy change costs | A new shot | An audio regen plus a re-sync | An audio regen plus a re-render |
Note what is not on that table: a single quality winner. Route A has the highest ceiling because nothing was composited. Route B has the highest floor because the picture was already real. Route C is a function of its input.
The copy-change test
Here is the scenario that decides it. Your 30-second spot is approved. The day before launch, compliance changes "up to 40%" to "up to 35%". Six words move. Nothing else about the video changes.
Route A. You regenerate. A new generation is a new seed, so you get a new performance, plausibly a slightly different face, certainly different framing and timing. The shot that was approved no longer exists, and the version you now have has never been through review. You are not fixing a video, you are remaking one and re-approving it. Budget for several attempts before you get one as good as the one you lost, and note that on Veo 3.1 the job cost the catalog states runs from 80 credits up to 320 depending on settings, so "several attempts" is a real line item.
Route B. You regenerate the audio, which is billed per 1,000 characters and therefore costs a rounding fraction of a video render at this length, then re-run the lipsync pass over the identical footage. The picture that was approved is the picture that ships. Nothing about framing, performance or grade changed, because nothing about the footage changed. On VEED Lipsync the catalog puts a job in the 4 to 81 credit range depending on length; Sync Lipsync 2.0 runs higher for the precision.
Route C. Same shape as B, provided you kept the photo and the voice selection. Re-render from the same still with new audio. The face is the same face, because the face is a file.
That is the argument, and it is not close. If the copy might change, and on commercial work the copy always changes, routes B and C are recoverable and route A is not.
The rule
- One-off, nobody will ever revise it, and you want the highest ceiling: route A. Hero shots, concept films, anything where the generation is the artefact.
- You already have footage of a real person: route B. Dubs, line fixes, localised versions, anything where the human was already filmed.
- You need a presenter who does not exist yet, at volume: route C. UGC-style spots, product explainers, weekly series where the same face has to come back.
- The copy is not final: B or C, always, even if A would look better. You will spend the visual difference on re-approvals.
- Script runs over about forty words: B or C. Route A's duration ladder decides this for you.
One practical note for routes B and C: the lipsync call exposes emotion, expression, talking_style and sync_mode alongside the image and audio. Those are the parameters that move output quality most, and they are the ones people leave at defaults. Set the sync mode deliberately before you blame the model for a length mismatch. The AI lipsync tool page covers what each parameter does.
FAQ
Which route gives the most convincing result?
On a single clip with a good prompt, route A, because sound and picture were produced together and nothing had to be reconciled. But "most convincing" is measured on the clip you keep, and route A produces the most clips you throw away, because you cannot steer it toward a specific person. Over a project, route B usually delivers the better finished piece for less effort.
Can I keep the same face across twenty videos on route A?
Not reliably from prompts alone. Text descriptions do not pin identity tightly enough, and a new generation is a new draw. If the same face has to recur, either fix it with a reference image and a reference-to-video mode, or move to route C where the face is a file you reuse. That is the whole reason avatar workflows exist.
Is the photo route worth it if I could just film someone?
If you have a willing person and a phone, filming and running route B gives you a better result for less money most of the time. Route C earns its place when there is no one to film, when the presenter needs to exist in four languages, or when you need forty variants next week and nobody wants to be on camera forty times. The talking-head model shortlist is the place to start once you have decided which of those applies.
How should I budget a 30-second talking head?
By route, because the shapes differ. Route A is several short generations stitched together, each charged in full. Routes B and C are one longer render plus a cheap audio pass, and both are billed by output duration. The worked figures are in what a 30-second talking-head video costs, and the model-by-model quality comparison is in lipsync options compared.