Guides

    6 published lipsync rows that take a driving audio track in 2026

    The catalog has 19 lipsync models and only 9 list audio-to-lipsync; six of those nine have a /models page: Wan 2.2 Speech to Video, Kling Avatar Pro, LTX 2 Audio to Video, LTX 2.3 Audio to Video, HeyGen Image to Video, and VEED Fabric 1.0 — a script-only talking-head row is a different category.

    Versely Team6 min read

    The catalog has 19 lipsync models and only 9 list audio-to-lipsync; six of those nine have a /models page: Wan 2.2 Speech to Video, Kling Avatar Pro, LTX 2 Audio to Video, LTX 2.3 Audio to Video, HeyGen Image to Video, and VEED Fabric 1.0 — a script-only talking-head row is a different category.

    Not all 19 lipsync models take a voice track you supply is the four-bucket map: nine audio-to-lipsync, five video-to-lipsync, four premade avatars, one text-driven. This page is only the six published /models cards in the audio-to-lipsync nine. Bookmark these slugs. Do not twin nine “audio-to-lipsync” posts.

    A driving track is the category

    Audio-to-lipsync means you already have speech as a file. The mouth follows the WAV (or the voice you attach through the product). You do not type the monologue into a text-to-video prompt and hope the model hears it. If you do not have a track yet, you are not in this six. Record it, clone it, or run a TTS row first.

    Many of these six also list image-to-lipsync: a still you keep, plus audio. Lock the presenter’s still, drive the face from the track. Changing the still later resets recognition. The audio is the performance. The model is not inventing the words.

    AI lipsync is the launcher. Make a talking avatar video is the named agent phrasing. Cost is priced per model, shown before you confirm. Name the row. A bare “make it talk” is how you land on a script-only twin. Length usually follows the track. Most of these six publish empty duration lists because the WAV is the clock. Cut the line you will ship before you spend the floor on throat-clearing.

    The six published /models pages

    Do not “upgrade in the prompt.” Change the row.

    1. Wan 2.2 Speech to Video — Audio-to-lipsync and image-to-lipsync. Image required. Audio true. 720p ceiling. 50 credits. Static image plus audio; length follows the track. Social close-up, not a 4K master.

    2. Kling Avatar Pro — Audio-to-lipsync and image-to-lipsync. Image required. Audio true. 4K. 58 credits display, per-second billing (min 58, max 1380). Bring a plate that can hold 4K. Sketch the still and the line first.

    3. LTX 2 Audio to Video — Audio-to-lipsync only. Audio true. 20 credits, flat. requires_image is false. Features include audio visualization and music video. The track is the brief. Bring a plate anyway when identity matters.

    4. LTX 2.3 Audio to Video — Audio-to-lipsync. Audio true. 20 credits. 1080p. Aspects: auto, 16:9, 9:16. Start frame optional. Write performance around the uploaded track. Do not type the VO into the prompt and upload a different take.

    5. HeyGen Image to Video — Audio-to-lipsync. Audio true. Image required. 50 credits floor; per-second at 10 credits, max 1200. Upload a selfie, choose a voice. The kitchen behind them is not this row’s job.

    6. VEED Fabric 1.0 — Audio-to-lipsync and image-to-lipsync. Image required. Audio true. 720p. 40 credits. Image URL plus audio URL. Catalog copy also says “professional fabric and texture video generation” — read that next to the categories. This is a talking still, not a set build.

    Script-only talking-head is a different category

    A script-only talking-head row is text-to-lipsync or premade-avatars. You type words. The product invents or selects a voice. You did not supply a driving track.

    VEED Fabric 1.0 Text is the hybrid sibling: a photo plus a script, voice auto-generated. It is not “upload WAV.” Avatar X Text to Video is the published text-driven card: type a script, a stock avatar performs it. Premade libraries and video-to-lipsync redubs (footage you already shot) are further doors. If you only have a paragraph, you are not in audio-to-lipsync yet. Make the track. Then pick one of the six.

    This page is the published audio-in index, not a ranked twin.

    Name the files, then the row

    Hold the still (if the row wants one). Play the audio with your eyes closed. If either file would embarrass you, do not combine them.

    You have Published row Sticker
    Still + track, 720p close-up Wan 2.2 Speech to Video 50cr, image required
    Still + track, 4K mouth Kling Avatar Pro 58cr floor, image required
    Track first, plate optional, flat 20 LTX 2 Audio to Video 20cr flat
    Track first, 1080p, optional start frame LTX 2.3 Audio to Video 20cr, 1080p
    Selfie + chosen voice HeyGen Image to Video 50cr floor
    Still + WAV, 720p talking still VEED Fabric 1.0 40cr, image required

    The three audio-to-lipsync rows without a /models page stay off this index. Do not invent their names. Live catalog is scripts/_published-models.json. Caption anyway. Health and finance stills do not get a photoreal talking presenter on this pass — the YPP third bucket is the format, not the slug.

    FAQ

    Do all 19 lipsync models take a WAV I supply?

    No. Nine list audio-to-lipsync. Six of those nine have a /models page — the list on this post. Five want footage. Four are premade avatars. Text-driven talking-head is a different category. The 19-way split is the bucket map.

    Can I type the script into Wan 2.2 or Fabric instead of uploading audio?

    No. Those published rows are audio-to-lipsync (Fabric also image-to-lipsync). They hear a file. Script-only is Fabric 1.0 Text or a premade twin. Make the track, then come back.

    Which of the six requires a still?

    Wan 2.2 Speech to Video, Kling Avatar Pro, HeyGen Image to Video, and VEED Fabric 1.0 require an image. LTX 2 Audio to Video and LTX 2.3 Audio to Video do not gate on a still; bring one anyway when the face has to be someone you already approved.

    Is make-a-talking-avatar-video the same as picking one of these six?

    Make a talking avatar video is the in-app job for a face image plus audio. It still has to land on a row. Name Wan, Kling Avatar Pro, HeyGen, Fabric, or an LTX audio-to-video slug if matching matters. A script with no WAV is a different job.