The formula
credits = rate_per_second × seconds- rate_per_second
- set entirely by the model, and on some models by the resolution band too — 1 to 20 credits across the catalogue. It does not move with script length, voice or language.
- seconds
- the length of the finished clip, which for lipsync is the length of the audio you feed it. Trim the audio and the bill falls in exact proportion.
Worked example — A 30-second piece to camera on Sync Lipsync 2.0
- 1.rate_per_second = 5 credits
- 2.seconds = 30
- 3.5 × 30
= 150 credits
Why the same 30 seconds has a 20-fold price range
Linear per-second billing is the easiest formula on Versely to predict and the easiest one to overpay on. There is no entry fee, no minimum block and no rounding: one second costs the rate, sixty seconds cost sixty times the rate. That makes the arithmetic trivial and makes the model choice the entire decision.
The spread is what people miss. Every rate in the table below animates a face to speak, and every one is billed on the same rate-times-seconds shape — but the rates differ by 20×. On a single 30-second clip that is 30 credits against 600. Across five clips a week for a month it is 600 against 12,000.
The same shape prices video extension and motion-control jobs, which is why a long extend can quietly become the largest single charge in an account: nothing in the formula caps it except the model's own maximum duration.
The lipsync rate card, second by second
Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.
Every per-second rate in the lipsync catalogue
One representative model per distinct rate. A rate table, not a ranking — cheaper is not better, and the buyer's guides are where models get judged.
| Credits per second | Example model | 10s | 30s | 60s |
|---|---|---|---|---|
| 1 | VEED Avatar Audio | 10 | 30 | 60 |
| 4 | HeyGen Avatar V3 | 40 | 120 | 240 |
| 5 | Sync Lipsync 2.0 | 50 | 150 | 300 |
| 6 | Kling Avatar Standard | 60 | 180 | 360 |
| 8 | VEED Fabric 1.0 (SD) | 80 | 240 | 480 |
| 9 | Sync Lipsync 2.0 Pro | 90 | 270 | 540 |
| 10 | Wan 2.2 Speech to Video (SD) | 100 | 300 | 600 |
| 12 | Kling Avatar Pro | 120 | 360 | 720 |
| 15 | VEED Fabric 1.0 (HD/4K) | 150 | 450 | 900 |
| 17 | Sync React 1 | 170 | 510 | 1,020 |
| 20 | Wan 2.2 Speech to Video (HD/4K) | 200 | 600 | 1,200 |
Every cell is the rate multiplied by the duration — that is the entire formula. 22 rate rows across 17 models collapse into these 11 distinct prices.
Where each model stops
The catalogue publishes a floor and a ceiling per model. The ceiling is the rate times the longest clip it will render, so it doubles as a maximum-length check.
| Model | Published credit range | Longest clip that implies |
|---|---|---|
| VEED Avatar Audio | 3 – 60 | 60s at its top rate |
| VEED Avatars | 3 – 70 | 70s at its top rate |
| VEED Lipsync | 4 – 81 | 81s at its top rate |
| LTX 2.3 Audio to Video | 50 – 200 | 20s at its top rate |
| Sync React 1 | 84 – 251 | 15s at its top rate |
| HeyGen Avatar V3 | 17 – 408 | 102s at its top rate |
| Sync Lipsync 2.0 | 25 – 600 | 120s at its top rate |
| Kling Avatar Standard | 29 – 675 | 113s at its top rate |
| Sync Lipsync 2.0 Pro | 42 – 1,000 | 111s at its top rate |
| HeyGen Avatar V5 | 50 – 1,200 | 120s at its top rate |
| HeyGen Image to Video | 50 – 1,200 | 120s at its top rate |
| Kling Avatar Pro | 58 – 1,380 | 115s at its top rate |
| VEED Fabric 1.0 | 40 – 1,800 | 120s at its top rate |
| VEED Fabric 1.0 Text | 40 – 1,800 | 120s at its top rate |
| Wan 2.2 Speech to Video | 50 – 2,400 | 120s at its top rate |
| Wan 2.2 Speech Turbo | 50 – 2,400 | 120s at its top rate |
| VEED Fabric 1.0 Fast | 50 – 2,400 | 120s at its top rate |
What changes a lipsync bill
Which model renders it
The only lever that matters. Same 30 seconds, 30 credits or 600.
Audio length
Directly proportional. Cutting a rambling 45-second read to a tight 30 removes exactly a third of the bill.
Resolution, on models that price it
Several lipsync models publish separate SD and HD/4K rates, and the gap between them is often as large as the gap between two different models.
Retakes
Every retake is a fresh render at the full rate. There is no partial charge for a clip you throw away and no discount on the second attempt.
What a lipsync charge does not cover
- The voiceover itself — if you are generating the audio rather than recording it, that is billed separately per 1,000 characters of script.
- Captions burned onto the finished clip, which are their own render with their own preset multiplier.
- The face image, if you are generating that rather than uploading a photo.
How many talking heads a month of credits buys
A Monthly plan's 580 credits covers 19 thirty-second clips at the 1-credit rate, or 0 at the 20-credit rate. Same plan, same clip, 19 clips of difference.
Plan sizes and what each grants are on the cost hub.
Lipsync models on the per-second meter
Every model page shows its own price matrix, its published credit range, and the option table this formula reads from.
VEED Avatar Audio
per_secondVEED · 3–60 credits
HeyGen Avatar V3
per_secondHeyGen · 17–408 credits
Sync Lipsync 2.0
per_secondSync · 25–600 credits
Kling Avatar Standard
per_secondKling · 29–675 credits
VEED Fabric 1.0
resolution_based_per_secondVEED · 40–1800 credits
Sync Lipsync 2.0 Pro
per_secondSync · 42–1000 credits
Wan 2.2 Speech to Video
resolution_based_per_secondWan · 50–2400 credits
Kling Avatar Pro
per_secondKling · 58–1380 credits
Gemini Omni Flash Image to Video
per_secondGoogle · 50–125 credits
Gemini Omni Flash Text to Video
per_secondGoogle · 50–125 credits
Gemini Omni Flash Edit
per_secondGoogle · 63–1500 credits
Gemini Omni Flash Reference to Video
per_secondGoogle · 50–125 credits
Frequently asked questions
Is a 60-second talking head exactly twice the price of a 30-second one?+
Yes. Per-second billing is linear with no entry fee and no minimum block, so doubling the duration doubles the credits precisely. That is not true of the block-priced or base-plus-marginal models used for generated video.
Why is one lipsync model 20 times the price of another?+
Because the rate is a property of the model, not of your clip. The catalogue publishes per-second lipsync rates from 1 to 20 credits, and every one of them is billed on the same rate-times-seconds formula.
Does the language of the script change the price?+
No. Nothing in this formula reads the language — only the duration. A 30-second clip costs the same in Japanese as in English. Translating a finished video into other languages is a different job on a different formula.
What is the longest clip I can render?+
Each model publishes its own ceiling. The second table converts every published maximum back into seconds at that model's top rate, so the length limits are readable at a glance.
Do I pay for a render that comes out wrong?+
A completed render is charged whether or not you like it. The cheapest insurance is a short test on the same model first — at 1 to 20 credits a second, a five-second test is a fraction of the full job.
Different job, different arithmetic
These are not variations on the formula above — each one meters something else entirely.
What does exporting in 4K instead of 1080p cost?
credits = rate_per_second(resolution) × secondsIt switches the per-second rate by resolution band, and the premium differs by model.
What does a clip longer than five seconds cost?
credits = base_first_5s + rate_per_extra_second × (seconds − 5)It charges an indivisible entry block, then meters only the seconds beyond it.
What does a finished 30-second AI ad cost end to end?
credits = (image_price × shots) + (video_rate × seconds) + ceil(chars ÷ 1,000 × 12) + (lipsync_rate × seconds) + exportIt sums four unrelated meters, so no single rate predicts the total.
Which model should I actually pick?
This page prices the job. The buyer's guides rank the models, with the sort criterion stated on every one.
Best AI lipsync model
Filtered to lipsync models, sorted by best leaderboard position (unranked models after, cheapest complete job first).
Cheapest AI lipsync model
Filtered to lipsync models that publish a complete-job cost, sorted by the credits one job costs, ascending.
Best AI model for talking head videos
Filtered to lipsync models with a premade-avatars, text-to-lipsync or image-to-lipsync category, sorted by the credits one complete job costs, ascending.
Jobs billed this way
Further reading
See the price before you spend it
Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.