Guides

    A finished 30-second ad is four different meters added together

    Still, motion, VO and lipsync bill as flat-per-call, per-second, per-1,000-characters and per-second-again; estimating from any one of them is how quotes go wrong.

    Versely Team6 min read

    Still, motion, VO and lipsync bill as flat-per-call, per-second, per-1,000-characters and per-second-again; estimating from any one of them is how quotes go wrong.

    Every other page on /cost isolates one formula. A finished 30-second AI ad is what happens when a real deliverable touches four of them at once. The stages do not contribute equally. Halving the script barely moves the total. Changing the video model changes everything. A quote that started from “lipsync is n credits a second” is missing three other meters and the export that sits on top.

    Four meters that do not share a unit

    The stacked formula on that scenario is:

    credits = (image_price × shots) + (video_rate × seconds) + ceil(chars ÷ 1,000 × 12) + (lipsync_rate × seconds) + export

    Those terms are different objects:

    • Reference stills — flat per call. The cheapest image models in the catalogue sit at 1 credit a generation; the price point most of them share is 4. Aspect ratio and prompt length do not move it.
    • Footage — per second, at the model’s resolution band. Lean: LTX 2.3 Text to Video Fast at 4 credits a second (SD/HD). Premium: MiniMax H3 at 26. Same 30 seconds, 120 or 780 credits.
    • Voiceover — per 1,000 characters, rounded up, at 12 credits on the rate most speech models publish. A 600-character read is 8 credits. Runtime and voice do not move it. Voiceover for a script is that meter alone.
    • Lipsync — per second again, on a separate rate card. Lean: VEED Lipsync at 1 credit a second. Premium: HeyGen Avatar V5 at 10. A 30-second talking head isolates this line.

    Export is one editor charge for the finished cut, with free 480p previews before it. Captions are their own render with a 1× or 2× preset multiplier. Neither is inside the four-stage subtotal. You cannot convert per-second into per-1,000-characters. Adding the meters is the quote.

    The lean build versus the premium build

    Same brief, same 600-character script, same 30 seconds. Only the model at each stage changes.

    Lean (cheapest credible row per stage): 1 credit still + 120 footage + 8 voiceover + 30 lipsync = 159 credits, plus captions and one export.

    Premium: 4 + 780 + 8 + 300 = 1,092 credits, plus the same extras. That is about 6.9× the lean build from model selection alone.

    Where the money actually goes, as a share of each total:

    Stage Lean Premium
    Footage ~75% ~71%
    Lipsync ~19% ~27%
    Voiceover ~5% under 1%
    Reference still under 1% under 1%

    The script did not change. Its share collapsed because the two per-second cards got steeper. Footage is the largest line in both builds. Lipsync is second, and independent — premium footage can still lipsync cheaply. A Monthly plan’s 580 credits covers three lean 30-second ads, or half of one premium ad, before retakes.

    Optimise the stage that dominates

    The cheapest single change is dropping the footage model a tier. It is a per-second rate, so the saving compounds across the whole runtime. 4 against 26 credits a second is the single biggest number on the stacked page.

    Runtime multiplies two of the four stages at once. A fifteen-second cut is not half an ad’s cost — the still and the VO do not halve — but it is close to half of the two per-second lines. If the brief can live at 15s, that is a real lever. If it cannot, do not pretend a shorter script is the same lever.

    Script length barely moves the total. At 12 credits per 1,000 characters the voiceover is a few percent of even the lean build. Write it the length the ad needs.

    Retakes at the footage stage are the expensive habit. Storyboard on cheap stills, preview, then generate expensive seconds once. Editor previews and final export exist so a 480p pass is free until the cut is locked. Regenerating thirty seconds of premium footage to “see the type” is how 1,092 credits becomes two ads.

    What the stage totals leave out

    The 159 / 1,092 figures assume each stage runs once. They do not include:

    • Captions, billed as their own render with a 1× or 2× preset multiplier.
    • The final editor export — one charge, with free previews before it.
    • Dubbing the finished ad for other markets, a flat per-job charge per language.
    • Anything you regenerate.

    A talking-head quote from /cost/30-second-talking-head-video is the lipsync line only: rate × seconds, script and language ignored. A VO quote from /cost/voiceover-for-a-script is the character line only. The AI lipsync surface is how you run that class of model, not a fifth formula. Add the four meters when the deliverable is an ad. Stop estimating from the one you opened first.

    FAQ

    Why is the range between the two builds so wide?

    Two of the four stages are per-second meters and both rate cards vary enormously. Footage runs 4 to 26 credits a second in this worked example and lipsync 1 to 10, and each multiplies by 30 seconds.

    Should I shorten the script to save credits?

    Almost never. At 12 credits per 1,000 characters the voiceover is a few percent of the total. Write the script the length the ad needs.

    Does 159 or 1,092 include captions and the export?

    No. Both sit on top and both are quoted at render time. Captions carry a 1× or 2× preset multiplier; the export is a single charge with free 480p previews before it.

    Can I quote from the lipsync rate alone?

    Not for a finished ad. Lipsync is one of four meters, and in the lean build it is about a fifth of the subtotal. The footage line is the one that dominates. Use the stacked scenario, not a single hub row.