The formula
credits = ceil( characters ÷ 1,000 × 12 )- characters
- the literal length of the text you submit, spaces and punctuation included. Not words, not seconds, not sentences.
- 12
- credits per 1,000 characters — the rate 16 of the 17 speech models in the catalogue publish, and the same figure the app's own estimator returns before it charges.
- ceil
- the total rounds up to a whole credit, which is why the shortest possible script still costs something.
Worked example — A 90-word ad read — 520 characters including spaces
- 1.520 ÷ 1,000 = 0.52
- 2.0.52 × 12 = 6.24
- 3.rounded up to a whole credit
= 7 credits
The one meter that reads your input, not the output
Text-to-speech is the only formula on Versely that meters your input rather than its output. Everything else on this hub counts what came back — seconds of video, megapixels of image, renders exported. Speech counts what you sent.
That inversion is what makes voiceover cheap and what makes it easy to misjudge. A slow, dramatic read of 500 characters produces far more audio than a fast one, and both cost the same. Adding a two-second pause is free. Switching to a language whose translation runs longer on the page costs more, because the character count rose.
It also means you can price a script before you have heard a word of it. Count the characters, divide by 1,000, multiply by 12, round up. That is the whole calculation, and it does not change with the voice you choose.
Script length, rounding and rate uniformity
Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.
Script length to credits
Character counts include spaces and punctuation. The 1,001-character row is there on purpose — it shows the rounding.
| Characters | Raw arithmetic | Credits charged |
|---|---|---|
| 250 | 0.250 × 12 = 3.00 | 3 |
| 500 | 0.500 × 12 = 6.00 | 6 |
| 750 | 0.750 × 12 = 9.00 | 9 |
| 1,000 | 1.000 × 12 = 12.00 | 12 |
| 1,001 | 1.001 × 12 = 12.01 | 13 |
| 1,500 | 1.500 × 12 = 18.00 | 18 |
| 3,000 | 3.000 × 12 = 36.00 | 36 |
| 6,000 | 6.000 × 12 = 72.00 | 72 |
| 12,000 | 12.000 × 12 = 144.00 | 144 |
One character past a thousand costs a whole extra credit, and nothing more until the next thousand. Trimming a script to land just under a boundary is worth a credit; agonising over it is not.
The same content at five script lengths
Assumes roughly five characters per word including the space after it — check your own script rather than trusting the average.
| What you are writing | Approx. words | Approx. characters | Credits |
|---|---|---|---|
| A hook line for a Reel | 20 | 100 | 2 |
| A 30-second ad read | 75 | 375 | 5 |
| A 60-second explainer | 150 | 750 | 9 |
| A five-minute narration | 750 | 3,750 | 45 |
| A twenty-minute podcast script | 3,000 | 15,000 | 180 |
Word counts are your input, not a promise about runtime — how long the finished audio runs depends on the voice and the delivery, and neither is metered.
How uniform the speech rate is
Unlike video, speech is priced almost identically across the catalogue — which model you use is a voice decision, not a budget one.
| Credits per 1,000 characters | Models at that rate |
|---|---|
| 6 | 1 of 17 |
| 12 | 16 of 17 |
17 speech models, 2 distinct rates. Compare that with lipsync, where one job carries a 20× spread.
What changes a voiceover bill
Cutting words
Every 1,000 characters you delete is 12 credits back. Editing the script is the only real optimisation.
Translated versions
A script that renders longer on the page in another language costs more, purely because the character count rose. Nothing about the language itself is priced.
Voice, emotion and style settings
Free. They change the delivery, not the character count.
Number of speakers
Multi-speaker dialogue is billed on script length and speaker count together, so a two-hander is not the same script at the same price.
Rewrites
A regenerated take resubmits the whole script and is charged again in full. Settle the words before you generate the audio.
What the speech charge does not cover
- Attaching the finished audio to a video, which is a separate flat fee that does not scale with length.
- Cloning a voice from a sample — cloning is billed once per sample, after which the clone generates at this same rate.
- Lipsyncing a face to the audio you just generated, which is billed per second of clip.
How many words a month of credits speaks
A Monthly plan's 580 credits is about 48,333 characters of speech — roughly 9,666 words, or 64 sixty-second explainer scripts.
Plan sizes and what each grants are on the cost hub.
Speech models on the per-1,000-character meter
Every model page shows its own price matrix, its published credit range, and the option table this formula reads from.
Gemini 3.1 Flash TTS
per_kcharsSeed Audio 1.0
per_kcharsByteDance
Cartesia Sonic 3.5
per_kcharsCartesia
Cartesia Sonic 3
per_kcharsCartesia
Grok TTS
per_kcharsGrok
Inworld TTS 1.5 Max
per_kcharsInworld
Inworld TTS 2
per_kcharsInworld
Inworld TTS
per_kcharsInworld
Inworld Voice Clone
per_kcharsInworld
ElevenLabs Multilingual
per_kcharsKIE
ElevenLabs Speech Turbo
per_kcharsKIE
Qwen 3 TTS 0.6B
per_kcharsQwen
Frequently asked questions
Why do speech models show a headline price lower than 12 credits?+
Because the headline figure on a model card is not measured in this formula's unit. Speech is billed per 1,000 characters at 12 credits, and that is the number that governs the charge — a headline like Qwen 3 TTS Voice Design's 3 credits is a catalogue figure, not a rate you can multiply. Use the character count and the app's own estimate.
Am I charged for seconds of audio or characters of text?+
Characters of text. This is the one place on Versely where the meter reads your input. Two takes of the same script at different speaking speeds produce different runtimes and identical bills.
Do spaces and punctuation count?+
Yes — the count is the raw length of the string you submit, which is why every table here quotes characters rather than words.
Is it cheaper to generate one long file or several short ones?+
Several short ones cost slightly more, because each generation rounds up on its own. Four 300-character segments are 16 credits; the same 1,200 characters in one pass is 15.
Does a premium voice model cost more per script?+
Barely, if at all — 16 of the 17 speech models bill the same 12 credits per 1,000 characters. Choose the voice you want rather than the one you assume is cheapest.
Different job, different arithmetic
These are not variations on the formula above — each one meters something else entirely.
What does a 30-second AI talking head cost?
credits = rate_per_second × secondsIt bills a flat rate for every second of output, so length is the only lever.
What does an AI video with sound cost?
credits = rate_per_second(audio on | audio off) × secondsIt flips the per-second rate with an audio toggle rather than adding a separate fee.
What does a finished 30-second AI ad cost end to end?
credits = (image_price × shots) + (video_rate × seconds) + ceil(chars ÷ 1,000 × 12) + (lipsync_rate × seconds) + exportIt sums four unrelated meters, so no single rate predicts the total.
Which model should I actually pick?
This page prices the job. The buyer's guides rank the models, with the sort criterion stated on every one.
Jobs billed this way
Further reading
See the price before you spend it
Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.