Not all 19 lipsync models take a voice track you supply
Nine are audio-to-lipsync, five take video, four are premade avatars, one is text-driven — so speech-in is not one route.
Nine are audio-to-lipsync, five take video, four are premade avatars, one is text-driven — so speech-in is not one route.
/without/showing-your-face is the routing document that refuses to call the lipsync class “audio-driven” wholesale. The class is 19 models that animate a face in time with speech. Only nine of them carry the audio-to-lipsync category — an audio track you supply is an explicit input. The other ten still move a mouth. They do not all start from a WAV you recorded. Picking “a lipsync model” and then discovering you needed footage, a stock avatar, or a typed script is how a generate sits in the wrong queue.
Nine are audio-to-lipsync
Those nine are the route people mean when they say “I have a voice track, animate this face.” Many of them also sit in image-to-lipsync: a still you keep, plus audio. That is the borrowed-face path on the without page: write and produce the audio first, lock the presenter’s still, drive the face from the track, check the first and last two seconds for drift.
Published rows in that bucket include Kling Avatar Pro, Wan 2.2 Speech to Video, HeyGen Image to Video, VEED Fabric 1.0, and LTX 2.3 Audio to Video. The still is the identity. Changing it later resets recognition. The audio is the performance. The model is not inventing the words.
If you do not have a voice track yet, you are not in this nine. Generate or record speech first — voiceover for a script if the meter is characters — then come back. Do not pick an audio-to-lipsync row and type a monologue into the prompt field hoping it will speak.
Five take video you already shot
video-to-lipsync is a different job: redub footage you already have. You keep the performer, the framing, and the lighting. The mouth changes to a new track. That is not “build a face out of nothing.” It is not the nine.
Published examples: Kling Lipsync (“syncs a speaker’s mouth in an existing video to new audio”), Sync Lipsync 2.0, Sync React 1, VEED Lipsync. /best/best-lipsync-model ranks the full lipsync pool and calls this out: several of them redub footage you already shot, which is usually the deciding factor rather than price.
If you have no footage, these five will not save you. If you have footage and you upload a still instead, you are on the wrong input class. The AI lipsync tool lists text, audio, and video paths as separate features because they are. Lipsync video in the editor is the generate_lipsync job (image + audio). Re-syncing an already-shot video’s mouth is a different path — dub with a video-native engine, or a video-to-lipsync row. Mixing those two sentences is how a still gets sent to a model that wanted an MP4.
Four premade avatars, one text-driven
Four rows are premade-avatars: a stock presenter you pick, not a face you brought. HeyGen Avatar V5 and VEED Avatars are the published end of that shelf. Fastest to ship. Weakest as a series identity unless you pick one avatar and never change it. The without page’s “own a cast” route exists because a fresh stock face every week is not a channel.
One row is text-driven: Avatar X Text to Video — type a script, a stock avatar performs it, voice included. Scripts on that card run 50–1,500 characters. You did not supply a voice track. You supplied words. VEED Fabric 1.0 Text is the related published hybrid (text-to-lipsync and image-to-lipsync): a photo plus a script, voice auto-generated. It is not “upload WAV.”
So “I will just lipsync it” is four different uploads:
| You have | Class | What you do not have to bring |
|---|---|---|
| A voice track + a face still | audio-to-lipsync (9) | Footage |
| Footage to redub | video-to-lipsync (5) | A new identity |
| Nothing but a presenter slot | premade-avatars (4) | Your face, maybe not even audio |
| A script | text-driven (1) | A recording |
Speech-in is not one route. The 19 is a content type. The four buckets are the actual doors.
Pick the route, then the row
/without/showing-your-face picks by whether anyone needs to be on screen, and whether that person must be the same next week. Then you pick a row that matches the files you actually have.
- Someone must talk to camera, it cannot be you, you have audio → the nine, still locked.
- You already shot it → the five.
- You need a body on screen this afternoon and you will not be in it → premade, then stop changing the avatar.
- You have a script and no session → the text-driven row, or Fabric Text if you also have a still.
Make a talking avatar is the in-app job for a face image plus audio. A 30-second talking head bills the lipsync meter per second once a row is named; it does not care which input class you needed, and it does not include the VO or the still. Name the input first. The rate is the second question. A cheap video-to-lipsync model is not a bargain if you never had a video.
Caption anyway. Synthetic delivery is harder to follow at low volume than a real one. That step is independent of which of the 19 you used.
FAQ
Do all 19 lipsync models take a WAV I supply?
No. Nine are explicitly audio-to-lipsync. Five want footage. Four are premade avatars. One is driven by text. “Lipsync” is the content type, not the upload.
Is video-to-lipsync the same as generate_lipsync on a photo?
No. generate_lipsync animates a face image to provided audio. Video-to-lipsync redubs a clip you already shot. /best/best-lipsync-model keeps both in the pool and still treats the input as the deciding factor.
Can I use a premade avatar for a weekly series?
Only if you pick one and keep it. A new stock face every episode resets recognition. The without page’s “own a cast” route is the series version of this problem.
If I only have a script, which class am I in?
Text-driven — Avatar X Text to Video — or a text-plus-still hybrid like VEED Fabric 1.0 Text. You are not in the nine until you have a voice track.