What you have to supply
Capability tells you what a model produces. This tells you what it wants first, which is the part that decides whether you can use it today. 4 require a starting image; anything not on that list runs from a text prompt alone.
| Model | You must supply | It will also take |
|---|---|---|
| LTX 2 Audio to Video | A text prompt only | — |
| VEED Fabric 1.0 | A starting image | — |
| HeyGen Image to Video | A starting image | — |
| LTX 2.3 Audio to Video | A text prompt only | — |
| Wan 2.2 Speech to Video | A starting image | — |
| Kling Avatar Pro | A starting image | — |
Read straight off each model's own record — requires_image, requires_video_input, accepts_video_input, max_reference and the per-mode reference configuration. Nothing is inferred from a model's name or description.
All 6 models
Cheapest complete job first. Every row links to that model's spec page — tier and mode variants share a page with the model they are a variant of, so this is products rather than SKUs.
| Model | Provider | Credits per job | Capabilities |
|---|---|---|---|
| LTX 2 Audio to Video | LTX | 20 credits | this only |
| VEED Fabric 1.0 | VEED | 40-1800 credits | also image-to-lipsync |
| HeyGen Image to Video | HeyGen | 50-1200 credits | this only |
| LTX 2.3 Audio to Video | LTX | 50-200 credits | this only |
| Wan 2.2 Speech to Video | Wan | 50-2400 credits | also image-to-lipsync |
| Kling Avatar Pro | Kling | 58-1380 credits | also image-to-lipsync |
How these models are billed
Three different billing shapes across the roster, which is why a straight comparison of headline figures between two of these models can mislead.
- Charged per second of output3 models
- Charged per second, at a rate that changes with output resolution2 models
- Charged a flat rate per generation1 model
All figures are Versely credits. A credit figure marked as a headline rate is not the cost of a complete job. Your plan's credit allowance is on the pricing page.
Looking for a ranked verdict?
This page is the complete list and the input requirements. It deliberately does not pick a winner. Versely already ranks a wider set that includes these models, and that verdict lives on its own page.
See the ranked verdictMore capability pages
Every model on Versely · Model providers · Ranked buyer guides
Frequently asked questions
How many audio-to-lipsync models does Versely have?+
6 with a spec page — all of them lipsync & avatars models, from 5 providers. Tier and mode variants of the same model share a page with the model they are a variant of, so this is products rather than SKUs.
What do audio-to-lipsync models need as input?+
4 of the 6 require a starting image, 0 require a video file, 0 accept a video without requiring one and 0 take reference images. Anything not on that list runs from a text prompt alone, and every figure here is read off the model's own record rather than inferred from its name.
What does an audio-to-lipsync job cost?+
Complete jobs run 20 to 58 credits across the 6.
Which audio-to-lipsync model is best?+
This page is the complete list and the input requirements — it deliberately does not pick a winner. Versely's ranked verdict over a wider set that includes these models lives at /best/best-lipsync-model.
Run every one of them on one subscription
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.