The formula
credits = (image_price × shots) + (video_rate × seconds) + ceil(chars ÷ 1,000 × 12) + (lipsync_rate × seconds) + export- image_price × shots
- flat per call — the reference stills and storyboard frames, priced per generation.
- video_rate × seconds
- per second of footage, at whatever rate the chosen model and resolution band set.
- ceil(chars ÷ 1,000 × 12)
- per 1,000 characters of script, rounded up. A 600-character read is 8 credits.
- lipsync_rate × seconds
- per second again, but on an entirely separate rate card from the video generation.
- export
- one editor export charge for the finished cut, with free 480p previews before it.
Worked example — The lean build, stage by stage
- 1.1 reference still at 1 credit
- 2.30s of footage on LTX 2.3 Text to Video Fast at 4/sec = 120
- 3.600-character voiceover = 8
- 4.30s lipsync on VEED Lipsync at 1/sec = 30
= 159 credits, plus captions and one export
Four meters, one deliverable, unequal shares
Every other page on this hub isolates one meter. This one is what happens when a real deliverable touches four of them at once, and the answer is that the stages do not contribute anything like equally.
In the lean build below, the footage is the largest line and the voiceover is nearly a rounding error. In the premium build the voiceover's share falls further still, because the video and lipsync stages have both moved onto much steeper rate cards. The script did not change; its share of the bill went from 5% to 0.7%.
That is the useful lesson for anyone budgeting a campaign: optimise the stage that dominates your particular stack, not the one that feels expensive. Halving the script saves almost nothing. Changing the video model changes everything.
Two builds of the same ad, stage by stage
Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.
Lean build — a 30-second ad
Cheapest credible model at each stage. Captions and the final export are quoted at render time and sit on top.
| Stage | Meter | Arithmetic | Credits |
|---|---|---|---|
| Reference still | flat per call | 1 × 1 | 1 |
| 30s of footage | per second | 4 × 30 | 120 |
| Voiceover | per 1,000 chars | ceil(600 ÷ 1,000 × 12) | 8 |
| 30s lipsync | per second | 1 × 30 | 30 |
| Total | — | — | 159 |
Premium build — the same 30 seconds
Same brief, same script, same runtime. Only the model at each stage changed.
| Stage | Meter | Arithmetic | Credits |
|---|---|---|---|
| Reference still | flat per call | 4 × 1 | 4 |
| 30s of footage | per second | 26 × 30 | 780 |
| Voiceover | per 1,000 chars | ceil(600 ÷ 1,000 × 12) | 8 |
| 30s lipsync | per second | 10 × 30 | 300 |
| Total | — | — | 1,092 |
6.9× the lean build, from model selection alone.
Where the money actually goes
Each stage's share of its build's total, so you can see which one is worth optimising.
| Stage | Share of lean build | Share of premium build |
|---|---|---|
| Footage | 75% | 71% |
| Lipsync | 19% | 27% |
| Voiceover | 5% | 0.7% |
| Reference still | 0.6% | 0.4% |
What changes the total most
The video model
The dominant term in both builds. 4 against 26 credits a second is the single biggest number on this page.
The lipsync model
The second biggest, and independent of the first — you can shoot premium footage and lipsync it cheaply, or the reverse.
Runtime
Multiplies two of the four stages at once. A fifteen-second cut is not half an ad's cost, but it is close to half of the two per-second lines.
Script length
Barely moves the total — the voiceover is 5% of even the lean build. Write it properly rather than writing it short.
Retakes at the footage stage
The most expensive place to iterate. Storyboard on cheap stills before generating expensive seconds.
What the stage totals leave out
- Captions, billed as their own render with a 1× or 2× preset multiplier.
- The final editor export — one charge, with free 480p previews before it.
- Dubbing the finished ad for other markets, a flat per-job charge per language.
- Anything you regenerate. Every total above assumes each stage runs once.
How many finished ads a month of credits buys
A Monthly plan's 580 credits covers 3 lean 30-second ads a month — or 0.5 premium ones.
Plan sizes and what each grants are on the cost hub.
Models used in these builds
Every model page shows its own price matrix, its published credit range, and the option table this formula reads from.
Luma UNI 1 Max
flatLuma · 1–1 credits
Nano Banana 2
flatGoogle · 4–4 credits
LTX 2.3 Text to Video Fast
resolution_based_per_secondLTX · 24–320 credits
Minimax H3 Text to Video
per_secondMiniMax · 130–390 credits
Qwen 3 TTS Voice Design
per_kcharsQwen
VEED Lipsync
per_secondVEED · 4–81 credits
HeyGen Avatar V5
per_secondHeyGen · 50–1200 credits
Gemini Omni Flash Image to Video
per_secondGoogle · 50–125 credits
Gemini Omni Flash Text to Video
per_secondGoogle · 50–125 credits
Gemini 3.1 Flash TTS
per_kcharsNano Banana 2 Lite
flatGoogle · 4–4 credits
Nano Banana Edit
flatGoogle · 2–2 credits
Frequently asked questions
Why is the range between the two builds so wide?+
Because two of the four stages are per-second meters and both rate cards vary enormously. Footage runs 4 to 26 credits a second here and lipsync 1 to 10, and each multiplies by 30 seconds.
What is the cheapest single change I can make?+
Drop the footage model a tier. It is the largest line in both builds and it is a per-second rate, so the saving compounds across the whole runtime.
Should I shorten the script to save credits?+
Almost never. At 12 credits per 1,000 characters the voiceover is a few per cent of the total. Write the script the length the ad needs.
Does this include captions and the export?+
No — both sit on top and both are quoted at render time. Captions carry a 1× or 2× preset multiplier; the export is a single charge with free previews before it.
How much should I budget for retakes?+
Plan the footage stage twice and everything else once. Iterating on stills and previews is cheap; regenerating thirty seconds of premium footage is the most expensive habit available.
Different job, different arithmetic
These are not variations on the formula above — each one meters something else entirely.
What does a 30-second AI talking head cost?
credits = rate_per_second × secondsIt bills a flat rate for every second of output, so length is the only lever.
What does exporting in 4K instead of 1080p cost?
credits = rate_per_second(resolution) × secondsIt switches the per-second rate by resolution band, and the premium differs by model.
What does editing a video cost if I keep changing it?
credits = (0 × preview_renders) + (export_price × final_exports)It charges per finished export, so iteration is free and the count is the whole bill.
Which model should I actually pick?
This page prices the job. The buyer's guides rank the models, with the sort criterion stated on every one.
Best value AI video model
Filtered to video generators that hold a leaderboard position and publish a complete-job cost, sorted by (leaderboard Elo ÷ credits per job) descending.
Best AI video generator
Filtered to video generators, sorted by best leaderboard position (unranked models after, cheapest complete job first).
Best AI model on every Versely leaderboard
Filtered to models holding a position on at least one Versely leaderboard, sorted by that position, ascending. The leaderboard each position was earned on is named in the last column.
Jobs billed this way
Further reading
AI Content Creation Cost in 2026: Real Budget Breakdown by Tier
What AI content creation actually costs in 2026: solo creator, brand, and agency budgets line by line. Hidden costs, regenerations, and traditional comparisons.
Cost per Creative: AI vs Agency vs In-House
A real cost-per-creative breakdown for 2026: AI production vs agency retainers vs in-house teams, including the hidden costs nobody invoices.
See the price before you spend it
Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.