Guides

    add_veed_captions transcribes speech; it does not write a headline

    Auto-captioning is one named job on audible speech. A fixed overlay, a transcript-only file, or captions on a silent clip spend this row on nothing.

    Versely Team3 min read

    add_veed_captions transcribes speech and burns styled, word-timed captions. It does not write a headline you typed. It does not listen to a silent b-roll plate. It does not dub. Give it a clip with no audible words and you have spent a caption pass on a file that had nothing to transcribe.

    Transcribe and caption my video is the named agent job. Speech in; styled, timed captions out. Required: a video_url and a preset. Default preset is glass if you do not name a vibe. BASIC presets (21 styles) cost the standard rate; DYNAMIC presets (9 premium styles) cost double. Transcription supports 165 language codes. The video must contain audible speech. That last sentence is the whole failure mode.

    You can preview a caption style on the first few seconds before the full render. Use that. Full-captioning a look you have not seen is how DYNAMIC double-rate becomes a surprise.

    What the job is

    A finished picture-and-soundtrack file. The tool transcribes what is actually on the tape and burns VEED subtitle presets — Glass, Whisper, Fusion, Glide, and the rest. It is not the word-level "dynamic captions" path, and it is not a static overlay. Offer it as a finishing step after a generate, not as a generate.

    The editing-task twin is add captions to a video. This capability is the agent name that routes to add_veed_captions.

    If you need a sidecar transcript without a burn-in, that is transcribe audio to text or get a video transcript. Different outputs. Different spends.

    What it is not

    • A line you wrote. Add a text overlay burns caption_text you supply. "WAIT FOR IT" is not a transcription. Running VEED captions on a music-only clip hoping the hook line appears will not typeset your copy.
    • Timed lines you already wrote. Add timed text overlays takes your timestamps. Nothing is transcribed.
    • A new voice. Captions do not re-voice, dub, or lipsync. Dub a video is a different engine.
    • QC of whether the generate matched the prompt. That is review_generation. Captions will happily subtitle a clip that ignored the brief.

    Using this job as a headline tool, or skipping it and regenerating the whole video "with captions in the prompt," both waste credits. Prompts do not typeset.

    The test

    Is there audible speech you want burned on, in a chosen preset, on a picture that is otherwise done?

    Yes: this job. Name the preset. Preview if you are unsure. Confirm the tier.

    If the words are yours and were never spoken, use overlay. If there is no picture yet, generate first. Caption last.

    FAQ

    What happens if the clip has music but no speech?

    The transcriber has nothing to do. You still spent the pass. Overlay the line you wanted, or generate speech and then caption.

    Can I caption in Spanish if the speaker is in English?

    You can request a language code. This job transcribes; it is not a dub. If you need a new voice in Spanish, dub first, then caption the dubbed file.

    Why preview at all if I already like Glass?

    DYNAMIC looks cost double. Preview is the cheap way to discover you actually wanted Whisper. The sample is the first seconds of your video, not a stock demo.

    Is this the same as asking the video model to "burn subtitles"?

    No. Scene models do not typeset a timed transcript. They sometimes paint illegible glyphs. Captioning is a finishing tool on a locked file.