Versely

    Captioning a month of clips: the credit math

    Prices a realistic month of short-form captioning in base renders, and shows why the clips you cull move the bill as much as the preset tier does.

    Versely Team8 min read

    Most people budget captions by publishing cadence: thirty posts a month, thirty caption renders. That number is almost always wrong, and it is wrong in the same direction every time — low. The charge attaches to the clip you captioned, not the clip you published, and in a normal month those two counts are nowhere near each other.

    Here is how to build a caption line you can actually defend at the end of the month.

    The unit is a base render, not a credit

    Captioning is billed as a render whose price is multiplied by the preset tier, and the per-clip figure is quoted in the app before the render runs. That makes "credits" a bad unit for planning, because the figure depends on the clip in front of you. It makes base renders a good one: one base render is one clip captioned on a 1× preset, and everything else in this post is a multiple of that.

    Two multipliers act on it:

    • The preset tier. The basic tier is 1×, the premium animated tier is 2×. Same transcription work, same timing pass, different charge. Basic tier versus dynamic tier covers what the multiplier buys and which one the tool reaches for when nobody specifies.
    • The clip count. This is the one that actually moves your month, and it is the one nobody plans.

    What captions cost across a month of clips works the tier multiplier out against publishing cadence. This post takes the other variable.

    A real month, four ways

    Take a team publishing 24 short clips a month. Behind those 24, they render around 40 candidates — some scenes get regenerated, some cuts get abandoned, a couple of concepts die after review. That ratio is unremarkable; a 60% keep rate on short-form is not a bad month.

    Now caption it four ways.

    Workflow Clips captioned Tier Base renders
    Caption every candidate, default preset 40 80
    Caption every candidate, basic preset 40 40
    Caption only what ships, default preset 24 48
    Caption only what ships, 18 basic + 6 premium 24 mixed 30

    Eighty base renders against thirty, for the same 24 published videos. Two levers did that, and they pull independently. Dropping every clip from the premium tier to the basic tier saves 40 renders on its own. Captioning only the clips that ship saves 32 on its own. Neither cancels the other, which is why the bottom row lands below what either change reaches alone — and why the clip-count lever is worth having even on a team that has already standardised on the basic tier.

    The reason the first row is so common is that captioning feels like part of making the clip. It is not. It is part of finishing the clip, and finishing should happen after the cull, not before it.

    The three charges people forget

    Captioning a clip you then cut. Captions are burned into the pixels at render time. Trimming a captioned clip afterwards does not refund the caption, and re-cutting it so the wording no longer matches means captioning it again from scratch. Every caption applied before the edit is locked is a coin flip on whether you pay for it twice.

    Fixing a typo after the fact. Same mechanism, more annoying. A single word change in the copy is a whole new captioning job at the same tier. Read the transcript, not the video, before you commit — it is faster and it catches the things a viewer will screenshot.

    Comparing styles by captioning full videos. Rendering three whole clips in three candidate looks to pick one is three full-price jobs to make an aesthetic decision. The caption tools can render just the opening seconds of your own footage in a candidate style instead, which is what previewing a caption style on five seconds of your own video is for. Two short previews beat three full renders and give you a better answer, because you are judging the style against your actual lighting and your actual speaking pace.

    The one structural saving

    There is a fourth lever that is not about clip counts at all. Captioning as a standalone job is one charge. Captioning inside a single editor assembly — where the same render stitches, trims, mixes the audio and burns the captions — resolves to one export.

    So a workflow that goes cut the video → attach the music → caption it bills three times for a result that a single assembly render bills once for. If your month involves any multi-clip edits, describing the finished video in one render rather than chaining single-purpose operations collapses three charges into one — a structural change, not a tier one, and it applies whichever tier you land on.

    While you are iterating, the editor's preview pass renders at 480p at no credit cost, with a short per-user cooldown between passes, so checking caption wording, placement and timing costs nothing. The final export is the charged render. Editor previews and the final export is the full version of that arithmetic — the short version is that the number of exports is the entire bill, and the previews exist so you only need one.

    The budget line you can defend

    Write it as a formula, not a number, and it survives a change in cadence:

    monthly caption renders
      = clips_shipped
      × (1 + rework_rate)
      × weighted_tier_multiplier
    

    For the team above: 24 clips shipped, a 15% rework rate for copy fixes, and a weighted tier multiplier of 1.25 because a quarter of the clips run premium. That is 24 × 1.15 × 1.25, or about 35 base renders a month. Multiply by whatever the app quotes per clip for your typical length and you have a line item with three named assumptions in it, each of which someone can argue with.

    The three assumptions are the point. "Thirty clips, thirty renders" cannot be argued with because it does not say anything. "24 shipped, 15% rework, a quarter on the premium tier" tells you exactly where to push when the number comes back too high — and the answer is nearly always the rework rate, followed by the share of clips captioned before the edit was locked.

    Two places to check the mechanics rather than the maths: add subtitles automatically is the transcription path for spoken audio, and the wider captions and on-screen text group covers the cases that are not transcription at all — fixed hook lines, CTAs and timed overlays, which are a different tool and a different charge.

    FAQ

    Should I caption before or after the final cut?

    After, in almost every case. Captions are burned in, so any change to the edit that changes what is said means a second captioning job. The exception is when the caption is the edit decision — if you are cutting to the rhythm of on-screen text, judge the timing on the editor's 480p preview pass, which carries no credit cost and a short per-user cooldown, then lock the cut and caption once for real.

    Does a longer video cost more to caption?

    The caption render is quoted per clip in the app before it runs, and the tier multiplier applies on top of whatever that figure is. The multiplier itself does not change with length. For budgeting purposes, treat "one clip captioned" as your unit and let the app resolve the per-clip figure.

    Is it cheaper to caption a batch of clips at once?

    There is no batch discount — ten clips is ten jobs. The saving available is in not captioning clips you will not publish, which is a workflow change rather than a pricing one. If several clips are going into one finished video, though, assembling and captioning them in a single editor render bills once instead of once per clip.

    What about translated captions?

    Captioning transcribes what is spoken; it does not translate into another language. Translating a video into a different language is a separate job on a separate meter, and it does not inherit the caption tier multiplier. Budget it on its own line rather than folding it into your caption count.