The formula
credits = caption_render × preset_multiplier × clips- caption_render
- the base charge for transcribing and burning captions onto one video, quoted per clip in the app before the render runs.
- preset_multiplier
- 1 for any of the 21 basic presets, 2 for any of the 9 premium animated ones. There is no third tier and no partial multiplier.
- clips
- captioning is per video, so a month of posting multiplies the whole thing by your publishing cadence.
Worked example — 30 clips a month on a premium preset
- 1.preset_multiplier = 2 (premium tier)
- 2.clips = 30
- 3.caption_render × 2 × 30
= the same credits as captioning 60 clips in a basic style
When a style choice is a pricing choice
Captioning is the only job on this hub where a purely aesthetic choice is a pricing choice. The work is identical across every preset — the same speech-to-text, the same timing pass, the same burn-in. What doubles is the charge, because the animated presets are sold as a premium tier.
On a single clip that is invisible. At 30 clips a month it is the second-largest recurring decision in a short-form workflow after the video model itself — and unlike the model choice, it changes nothing about the footage.
There is also an escape hatch most people skip: you can render a few seconds of a clip in a candidate style to compare looks before committing to a full-length render. Choosing a style by previewing a fragment costs a fraction of choosing it by captioning three whole videos and picking one.
The two preset tiers across a posting schedule
Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.
The two preset tiers
Tier assignment comes from the captioning tool itself. The look changes; the transcription does not.
| Tier | Presets available | Multiplier | Cost of 30 clips |
|---|---|---|---|
| Basic | 21 presets — plain, legible, static | 1× | 30 base renders |
| Premium (animated) | 9 presets — glass, glide, fusion, terminal and similar | 2× | 60 base renders |
The multiplier applies per clip, so it compounds with cadence rather than being a one-off styling fee.
The multiplier across publishing cadences
Expressed in base renders, so the arithmetic holds whatever the per-clip figure is in your account.
| Clips per month | All basic | All premium | Extra renders you are buying |
|---|---|---|---|
| 8 | 8 | 16 | +8 |
| 15 | 15 | 30 | +15 |
| 30 | 30 | 60 | +30 |
| 60 | 60 | 120 | +60 |
| 90 | 90 | 180 | +90 |
A daily poster on premium presets pays the same caption bill as a twice-daily poster on basic ones.
What changes a caption bill
Preset tier
A clean 2×, applied per clip, for a purely visual difference.
Previewing before committing
Rendering a few seconds in a candidate style is far cheaper than captioning whole videos to compare looks.
Mixing tiers deliberately
Premium on the hero posts, basic on the rest. The multiplier is per clip, so it need not be an account-wide decision.
Captioning once, cutting after
Captions are burned into the pixels, so a re-cut that changes the wording means a second captioning charge. Lock the edit first.
What a caption render does not cover
- Translating the captions — captioning transcribes what is spoken; translation is a separate job.
- The video's own generation cost.
- Static titles and text you write yourself, which are a standard editor render rather than a captioning job.
Budgeting captions by render, not by credit
Because the per-clip figure is quoted at render time, budget this one in renders rather than credits: 30 clips a month is 30 caption renders on the basic tier and 60 renders' worth on the premium tier.
Plan sizes and what each grants are on the cost hub.
Frequently asked questions
Do premium caption presets transcribe more accurately?+
No. The transcription and timing are the same work across every preset. The tier changes the visual treatment and the charge, not the accuracy.
Which presets sit in the 21-preset basic tier?+
The plain, static, legibility-first looks — simple, plain, corpo and similar. The premium tier is the animated set: glass, whisper, glide, fusion, terminal, handwritten and the backdrop variants.
Does a longer video cost more to caption?+
The caption render is quoted per clip in the app before it runs, and the preset multiplier applies on top of whatever that figure is. The multiplier itself never changes with length.
Can I preview a style without paying for a full render?+
Yes — the caption tools can render just the opening seconds of your video in a candidate style so you can compare looks before committing to the whole thing.
Are captions charged again if I re-export the video?+
Captions are burned into the pixels at render time, so a new export with different wording is a new captioning job. Settle the copy before the final pass.
Different job, different arithmetic
These are not variations on the formula above — each one meters something else entirely.
What does editing a video cost if I keep changing it?
credits = (0 × preview_renders) + (export_price × final_exports)It charges per finished export, so iteration is free and the count is the whole bill.
What does dubbing one video into 10 languages cost?
credits = dub_job_price × languagesIt bills one flat charge per job and ignores the target language entirely.
What does a finished 30-second AI ad cost end to end?
credits = (image_price × shots) + (video_rate × seconds) + ceil(chars ÷ 1,000 × 12) + (lipsync_rate × seconds) + exportIt sums four unrelated meters, so no single rate predicts the total.
Jobs billed this way
Further reading
See the price before you spend it
Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.