Capability

    Audio-to-lipsync AI models

    6 Versely models list audio-to-lipsync among their capabilities — all of them lipsync & avatars models, from 5 providers. What separates them is less what they do than what they need from you: 4 of them will not start without an image. The input table below is the part that actually decides which one you can use.

    6 models5 providers4 need a file to start

    What you have to supply

    Capability tells you what a model produces. This tells you what it wants first, which is the part that decides whether you can use it today. 4 require a starting image; anything not on that list runs from a text prompt alone.

    ModelYou must supplyIt will also take
    LTX 2 Audio to VideoA text prompt only
    VEED Fabric 1.0A starting image
    HeyGen Image to VideoA starting image
    LTX 2.3 Audio to VideoA text prompt only
    Wan 2.2 Speech to VideoA starting image
    Kling Avatar ProA starting image

    Read straight off each model's own record — requires_image, requires_video_input, accepts_video_input, max_reference and the per-mode reference configuration. Nothing is inferred from a model's name or description.

    All 6 models

    Cheapest complete job first. Every row links to that model's spec page — tier and mode variants share a page with the model they are a variant of, so this is products rather than SKUs.

    ModelProviderCredits per jobCapabilities
    LTX 2 Audio to VideoLTX20 creditsthis only
    VEED Fabric 1.0VEED40-1800 creditsalso image-to-lipsync
    HeyGen Image to VideoHeyGen50-1200 creditsthis only
    LTX 2.3 Audio to VideoLTX50-200 creditsthis only
    Wan 2.2 Speech to VideoWan50-2400 creditsalso image-to-lipsync
    Kling Avatar ProKling58-1380 creditsalso image-to-lipsync

    How these models are billed

    Three different billing shapes across the roster, which is why a straight comparison of headline figures between two of these models can mislead.

    • Charged per second of output3 models
    • Charged per second, at a rate that changes with output resolution2 models
    • Charged a flat rate per generation1 model

    All figures are Versely credits. A credit figure marked as a headline rate is not the cost of a complete job. Your plan's credit allowance is on the pricing page.

    Looking for a ranked verdict?

    This page is the complete list and the input requirements. It deliberately does not pick a winner. Versely already ranks a wider set that includes these models, and that verdict lives on its own page.

    See the ranked verdict

    More capability pages

    Every model on Versely · Model providers · Ranked buyer guides

    Frequently asked questions

    How many audio-to-lipsync models does Versely have?+

    6 with a spec page — all of them lipsync & avatars models, from 5 providers. Tier and mode variants of the same model share a page with the model they are a variant of, so this is products rather than SKUs.

    What do audio-to-lipsync models need as input?+

    4 of the 6 require a starting image, 0 require a video file, 0 accept a video without requiring one and 0 take reference images. Anything not on that list runs from a text prompt alone, and every figure here is read off the model's own record rather than inferred from its name.

    What does an audio-to-lipsync job cost?+

    Complete jobs run 20 to 58 credits across the 6.

    Which audio-to-lipsync model is best?+

    This page is the complete list and the input requirements — it deliberately does not pick a winner. Versely's ranked verdict over a wider set that includes these models lives at /best/best-lipsync-model.

    Run every one of them on one subscription

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.