Comparisons

    VEED Caption Presets: Basic Tier vs Dynamic Tier

    Thirty caption presets split across two credit tiers, one retired tool name still causing confusion, and a practical rule for when to pay double.

    Versely Team7 min read

    Ask for "dynamic captions" in Versely today and you'll get something, but it might not be the something you're picturing from an old screenshot or a six-month-old blog post. That confusion is worth clearing up before comparing anything else, because it changes which of two very different systems you think you're choosing between.

    The correction first: there is no separate dynamic-captions tool anymore

    Versely used to run its own in-house animated captioner, called add_dynamic_captions. It's retired — marked as such directly in the tool registry, with an explicit instruction not to call it. VEED's subtitle engine is the only captioning system live in the app now, reached through a single tool, add_veed_captions.

    The confusion this causes is specifically about naming, not capability: the old tool's name was "dynamic captions," and the current tool's premium tier is also called "DYNAMIC." Those are two unrelated things that happen to share a word. If you're working from anything written before the switch, or a habit formed on the old tool, the fix is simple — add_veed_captions is where every caption request goes now, and "DYNAMIC" inside it is a preset tier, not a separate product.

    One tool, thirty presets, two tiers

    With that cleared up, here's what add_veed_captions actually offers: thirty preset looks split into two tiers that differ in credit cost and in what kind of visual treatment they apply.

    BASIC — 1× credit cost, 21 presets. These are the everyday, legible styles: simple, plain, beans, corpo and others in the same register — built for readability first, with a straightforward look rather than a performed one. This is the tier for accessibility-first subtitling, long-form repurposed content, or anything where the caption's job is to be read accurately rather than noticed.

    DYNAMIC — 2× credit cost, 9 presets. Glass, whisper, glide2, fusion, glide, terminal, handwritten, backdrop and backdrop2. These are word-level animated, premium visual treatments — the kind built to double as a hook element in its own right, not just a transcript on screen. Glass is the tool's own default when nobody specifies a vibe, which is worth flagging on its own: ask for captions without naming a style and you're opted into the 2× tier by default, not the 1×.

    Both tiers transcribe against the same underlying speech recognition, covering 165 language codes — everything from en-US and en-GB through es-MX, ja-JP, ar-SA and dozens more. The tier choice changes how the words look. It has no bearing on which languages get heard correctly.

    What you're actually paying double for

    It's worth being precise about what the multiplier buys, because "premium" undersells the actual difference. A BASIC preset renders your transcript in a fixed style — a font, a color, a stroke, a static position — applied uniformly across every line. A DYNAMIC preset does something structurally different: word-level animation, where individual words highlight, glide, or reveal in sync with the audio rather than the whole line appearing and disappearing as a block. That's closer to a small motion-graphics pass than a font choice, which is a reasonable place for a credit multiplier to sit.

    The economics scale exactly the way you'd expect: a single clip, the 2× tier is a rounding error. Run it across a real publishing cadence and it stops being one. What captions cost across a month of clips works the multiplier out at volume — the short version is that the tier decision, repeated over a month of posting, is worth roughly as much as your video model choice, for a purely aesthetic pick that changes nothing about the underlying transcription work.

    A rule for picking, instead of guessing

    The honest answer to "which tier should I use" depends on what the caption is doing in the video, so here's the actual decision, not a style preference:

    • The caption is carrying part of the hook — it needs to grab attention on its own, independent of the footage. Reach for DYNAMIC. This is the one case where the word-level animation is doing real creative work, not just decoration.
    • The caption exists so the video is watchable on mute — an interview, a tutorial, a long-form repurpose. Reach for BASIC. Nobody's retention is improved by an animated word reveal on a talking-head explainer; legibility is the whole job.
    • You're not sure yet. Don't render the full clip in either tier to find out. preview_caption_style trims the first few seconds of your actual video and renders just that fragment in a candidate style — a cheap way to compare looks before committing a full render to one.

    Building the comparison in Versely

    A concrete way to run the choice rather than defaulting into it:

    1. Ask the agent to preview two or three candidate styles on your actual clip"Preview this video's first five seconds in the 'simple' basic preset and the 'glass' dynamic preset so I can compare." This calls preview_caption_style against short fragments instead of two full renders.
    2. Judge on the real footage, not a style name. A preset that reads well on a demo clip can behave differently against your actual lighting, font size at your platform's crop, and your speaking pace.
    3. Commit to one tier for the full render. "Add VEED captions to the full video using the glass preset, language en-US" calls add_veed_captions with your chosen preset, at whichever tier that preset belongs to.
    4. Hold the choice across a series. Caption styling that stays consistent across a run of videos does brand-recognition work in the feed before anyone reads a word — switching tiers clip to clip undercuts that more than either tier alone would cost you.

    If you specifically need speech turned into subtitles rather than text you already wrote burned onto the video, add subtitles automatically is the task page that walks the same add_veed_captions tool from the transcription side of the decision — useful background if the request in front of you isn't actually a caption at all, but a fixed headline or CTA that only looks like one.

    One more naming note, unrelated to captions

    VEED shows up as two separate things on Versely, and it's worth not conflating them while we're on the subject. Everything above is VEED's subtitle engine. VEED also supplies a separate roster of lipsync and avatar video models, ranked independently on best VEED model — a completely different product line, same company name. If a request in front of you is about a talking avatar rather than a caption style, that's the page to check, not this one.

    FAQ

    Is the old add_dynamic_captions tool still available for existing workflows? No. It's marked retired in the tool registry with an explicit instruction not to call it. Any workflow or saved prompt still referencing it needs to move to add_veed_captions, choosing a BASIC or DYNAMIC preset in place of whatever the old tool used to render.

    Can I switch a video from BASIC to DYNAMIC without re-transcribing? Yes — the transcription and word timing are the same underlying pass regardless of tier; only the visual treatment changes between presets. That's also why previewing a few seconds in each candidate style is cheap relative to committing to a full render twice.

    Does a longer video cost proportionally more to caption? The multiplier is about the preset tier, not the runtime — a BASIC preset is 1× and a DYNAMIC preset is 2× of whatever the base caption render costs for that clip. What captions cost across a month of clips breaks down how that base charge scales with a real publishing cadence.

    The takeaway

    Thirty presets, two tiers, one retired tool name that no longer applies. BASIC covers the 1× accessibility-and-legibility case in 21 flavors; DYNAMIC covers the 2× hook-carrying case in 9 — and the tool defaults to a DYNAMIC preset the moment nobody specifies otherwise, which is worth knowing before it quietly doubles a month of caption renders you assumed were running at the standard rate.