VEED burn-in waits until the spoken take is the keeper
add_veed_captions transcribes and burns styled subtitles. Run it on a miss, or before picture lock, and you pay for type on audio you will throw away.
add_veed_captions transcribes speech and burns styled subtitles into the pixels. It is a finishing pass. Run it on a take you are still regenerating, or before picture lock, and you pay for type on a soundtrack that will not ship.
The task page is add captions to a video. The engine is VEED. list_caption_fonts and preview_caption_style are how you pick a look. None of that is generation. None of it transfers to the next sample.
Captions are glued to this file's audio
The tool listens to this clip. It does not listen to the prompt. A new take has new timing, new breaths, new missed words. Karaoke highlights and cue breaks will not line up with the replacement. You cannot "move the captions over." You run the pass again.
That is the leak. BASIC presets bill at the 1× caption rate. The nine DYNAMIC looks — glass, glide, fusion, and the rest — bill at 2×. Preview a style on a few seconds if you are choosing a look. Do not burn a full DYNAMIC pass onto a candidate.
VEED is a transcription tool, not a translation tool, and not a headline tool. Language should match what was actually spoken. If you want a line nobody said — a hook, a CTA, a badge — that is add_video_captions, a different tool. If the audio is still a scratch, even a BASIC burn is early.
What has to be true first
Picture lock, in a pipeline that can always sample again, is a finishing contract: these shot ids, this timeline, this soundtrack. Picture lock when shots can still regen is the contract. For captions, lock is narrower and earlier:
- The take is the take. Nobody is about to "try it warmer."
- The voice is the voice you will ship, native or laid on later.
- Duration is the duration. You are not about to trim, extend, or punch a section out.
- You have already decided the words on screen are a transcript, not authored copy.
If any row is still open, skip the burn. Generate. Review. Attach or replace audio if the native track is not the mix. Then caption.
preview_caption_style exists so you can test glass versus a plain look on a short sample of the keeper. That sample is cheap compared with a full-length DYNAMIC render. Use it after you know which file is the file.
The order that does not leak
- Generate or shoot until the picture and the spoken line survive review.
- Mix or replace audio if the keeper still has scratch sound.
- Ask
list_caption_fontsif you need a real font id, preview one style on a few seconds, then burn.
Do not invert that because captions look like "progress." Styled type on a miss is a receipt for work you will delete. The AI caption generator is the same job in tool clothing. The rule does not change because the door does.
FAQ
Can I caption a first take and reuse the track on the regen?
No. Cues are timed to this file's speech. A new sample is a new clock. Burn again on the keeper, or do not burn yet.
Should I pick a DYNAMIC preset while I am still choosing a look?
No. Preview on a few seconds. DYNAMIC is 2× the BASIC rate for the whole duration. Pin a premium look after you know it reads on this plate, not while tasting.
Is this the tool for a headline I wrote myself?
No. add_veed_captions types what was said. A hook or CTA you authored is a text overlay. Mixing those jobs up is how you transcribe silence or ignore the line you actually needed.
Do I wait for a signed change-list?
You wait until this shot will not be regenerated for taste. A formal lock memo is nicer. A named keeper file is enough to spend the caption meter.