Timed text overlays are lines you time, not auto-captions
add_timestamped_captions burns copy you wrote at start_sec and end_sec. It does not transcribe speech.
Timed text overlays are lines you time, not auto-captions. Different text at different moments in the clip. add_timestamped_captions burns each line onto a standalone video at the start_sec and end_sec you set. Nothing is transcribed from the audio. Asking this job to "add captions" on a talking-head wastes credits on a typesetting pass that will not listen to a single word.
Add timed text overlays to my video is the named capability. You supply the lines and the timing. Position, size, and color apply to the set. The tool description is explicit: this is not a subtitle tool. For spoken-word subtitles use transcribe and caption my video (add_veed_captions).
What this job actually is
You wrote the copy. A listicle. A countdown. "Day 1" then "Day 30" then a punchline. Legal text that must not be invented by a model. You give each caption with its window. The agent burns the sequence. Flat per-video credit fee, shown before rendering.
The words are yours. The clock is yours. If you recut the picture after this pass, every timestamp is attached to the wrong frame. That is a finishing constraint, not a reason to skip picture lock.
A single persistent headline with no timing changes is add a text overlay to my video: one line, burned for the duration. Timed overlays exist because the text changes.
What wastes credits
Speech in, hope the tool hears it. Auto captions transcribe. Timed overlays do not. You will get silence on screen, or you will type a bad transcript by hand and pay a burn to display it. If the words are what someone said, transcribe.
One badge for the whole clip. "WAIT FOR IT" at the top for eight seconds is a fixed overlay. Splitting it into one timed line from 0 to 8 is the same graphic with extra fields.
Cards on a take you still might throw away. Kinetic type looks like an edit. It is labels on an edit. Burn once on a keeper. A second burn after you regenerate the shot is a second fee for copy that had to be rewritten to a new clock anyway.
An editor assembly that also needs music, trims, and captions. If the output is one finished video from several clips, build a video from clips, music, and captions is the one-pass render. Timed overlays on a file you are about to recut is the wrong order.
The test
Who wrote the words, and do they change on a clock you can name?
If you wrote them and you can name the seconds, this is the job. Lock the cut first.
If the words are speech, caption. If there is one line that never changes, overlay. If you do not have a cut yet, keep the copy in a doc. This is one named agent job. Using it for a different job wastes credits.
FAQ
Does this transcribe the video's audio?
No. You supply every line and every timestamp. There is no speech model on this path. add_veed_captions is the transcribe-and-burn job.
Can each line have its own font and color?
Font, size, color, and position apply to the whole set, not per line. If you need a different treatment on one card, that is a different overlay pass or a different tool, not a hidden per-line style in this job.
What if I need both subtitles and punch-in cards?
Two jobs. Transcribe the speech. Then, on the captioned keeper if you still want authored cards, burn the timed set. Do not ask timed overlays to do both.
Why not bake the text into the generate prompt?
Because then you cannot change a price, a spelling, or a language without a new generate. Timed overlays exist so the words can change without a new take. That is the credit argument, not a design preference.