Subtitles, Text Overlays and Timed Overlays Are Not the Same Tool
Ask for 'captions' and you might get a transcript of your voiceover, a supplied hook line, or a burned-in overlay that never touches the audio.
Ask for "captions" on a finished video and you can get three completely different results depending on what you actually meant: a word-for-word transcript of whatever was said, a single hook line burned across the whole clip, or a sequence of lines that change on their own schedule and were never spoken at all. All three get called "captions" in casual use. Only one of them is actually listening to the audio.
The naming trap, stated plainly
Versely's own tool names make this confusion almost inevitable if you don't already know the distinction: two of the four tools involved have the word "captions" directly in their function name, and neither one transcribes anything. add_video_captions burns a fixed text overlay you supply — a hook line, headline, badge, or CTA — and it is explicitly documented as not a subtitle tool. add_timestamped_captions burns a sequence of overlay lines you supply with your own timing, and it's documented the same way: nothing gets transcribed. The only tool of the four that actually listens to speech and generates text from it is add_veed_captions — the one whose name doesn't contain the word "subtitle" at all.
That's backwards from what the names suggest, and it's exactly why "just add captions" is an underspecified request. The fix isn't memorizing four tool names. It's one diagnostic question.
The one-line test
Before reaching for any of these, ask: does the text come from what was already said, or do I already have the exact words?
If the answer is "from what was said" — you have a voiceover or dialogue track and want it turned into readable on-screen text — you need actual transcription. That's the only job in this group that listens to audio at all.
If the answer is "I already have the words" — a hook line, a CTA, a headline you wrote yourself — transcription isn't just unnecessary, it's the wrong tool entirely, because there's no guarantee it would even produce the exact wording you need. From there it's a second, smaller question: is it one message for the whole clip, or does the text need to change at specific moments? One message is a fixed overlay. Changing text on a schedule is a timed sequence.
Subtitles: text that comes from the audio
add_veed_captions is the transcription tool — it runs the VEED Subtitles model against your video's audio and burns the result as styled, timed subtitles, synced to what was actually said. It's built for the case where you have spoken content and want a polished, social-ready caption track without typing out a transcript by hand.
The styling has real range: a BASIC tier covers 21 preset looks at the standard credit cost, and a DYNAMIC tier adds 9 premium, word-level animated presets — glass, whisper, glide, fusion and others — at double the credit cost of a basic preset. Transcription itself isn't limited to English; the tool supports 165 language codes, covering everything from en-US and en-GB through es-MX, fr-FR, ja-JP, ar-SA and dozens more. If your video has no spoken audio at all — a silent product montage, an ambient B-roll piece — this tool has nothing to transcribe, and reaching for it is the first sign the request was actually for an overlay instead.
Fixed overlay: one line for the whole clip
add_video_captions is what you want when the text is a single message that holds for the entire video and you already know exactly what it says — a hook line at the top of a UGC-style ad, a badge reading "NEW," a CTA that stays on screen throughout. You supply the caption text directly along with its position, and nothing about it is derived from the audio track. This is the tool most often reached for by mistake when someone says "add captions" but actually means "put my headline on this video" — the request has nothing to do with transcription, and using the transcription tool for it would produce the wrong text entirely, or nothing, if the clip has no dialogue to transcribe in the first place.
Timed overlays: text that changes, but wasn't spoken
add_timestamped_captions is the middle case: multiple lines of text you supply yourself, each with its own start_sec and end_sec, so the on-screen copy changes at specific moments through the clip — a beat-by-beat text story, a sequence of stat callouts, a series of short claims timed to cuts in the edit. Like the fixed-overlay tool, nothing here is transcribed; you're authoring both the words and their timing directly. The distinguishing question against subtitles is the same as above — if the lines are your own copy rather than a record of what was said on the audio track, this is the overlay tool, not the transcription one, no matter how closely the final look resembles styled subtitles.
The slideshow case: one overlay or N
add_text_overlay is the equivalent tool for a slideshow rather than a single video file, and it has its own small but important branch: send one overlay and the system auto-expands the same text and style across every image in the slideshow, or send a matching number of overlays — one per image — for different text on each slide. Getting this distinction backwards is its own common mistake: sending a single overlay when you meant per-slide captions repeats one line across a set that was supposed to tell a sequence, and sending a full array when a slideshow only needed one consistent line is unnecessary setup for the same result the auto-expand would have given for free.
Running the decision in Versely
A concrete pass through the diagnostic: for a UGC-style clip with real spoken dialogue that needs standard subtitles, the prompt is as simple as "Add styled captions to this video, transcribed from the speech, in the glass preset." That calls add_veed_captions with preset: "glass" against your video_url. For the same clip if what you actually want is a hook line layered on top — text that was never spoken — the request changes shape entirely: "Add the text 'Wait for it' as a bold overlay at the top of this video for the full clip" calls add_video_captions instead, with your exact wording supplied directly rather than inferred from anything on the audio track. If that hook needs to change three times through the clip rather than staying static, the same idea moves to add_timestamped_captions with three lines, each carrying its own start and end second. Naming which of the three you mean before you ask — audio-derived, fixed, or timed — is the entire fix for the naming trap above.
FAQ
If a tool has "captions" in its name, does that mean it transcribes speech?
Not necessarily, and this is the exact confusion worth avoiding. Two of Versely's caption-named tools — the fixed overlay and the timed overlay tools — explicitly do not transcribe anything; you supply the text yourself. Only the subtitle-specific tool actually listens to the audio and generates text from it.
What happens if I use the transcription tool on a video with no spoken audio?
There's nothing for it to transcribe, so it isn't the right tool for that job — a silent clip needs a supplied overlay instead, whether fixed for the whole video or timed to specific moments, since neither of those tools depends on audio content at all.
How do I add a hook line that wasn't part of my voiceover?
Use the fixed text overlay tool and supply your exact wording directly — this burns your specified text onto the video without attempting to transcribe or match it to the audio track, which is the correct approach whenever you already know precisely what the on-screen text should say.
Can captions and a separate text overlay both appear on the same video?
Yes — they're independent passes. Auto-transcribed subtitles from the spoken audio and a supplied hook line or CTA overlay can both be burned onto the same clip; they draw from different sources (audio versus your own supplied text) and don't conflict with each other.
What's the difference between sending one overlay and multiple overlays to a slideshow?
A single overlay auto-expands to apply the same text and style across every image in the slideshow. Sending one overlay per image instead gives each slide its own distinct text — the right choice whenever the slideshow is telling a sequence rather than repeating one consistent line.
Try the one-line test on your next clip: decide whether the text comes from the audio or from you, then run it through Versely's caption tool or the matching overlay tool — not whichever one happens to have "captions" in its name.