Auto-timed subtitles belong on the soundtrack you will ship
add_veed_captions times speech into burned-in subtitles. Doing that on a miss, or before picture lock, pays for a track you will recut.
Automatic subtitles are a comprehension pass: speech in, timed lines out, burned into the file. The engine is the same VEED path as styled social captions — add_veed_captions — pointed at a plain, readable preset instead of a hook look. It still has to wait. A miss has the wrong words. A recut has the wrong clock. Either way you paid for type that dies with the take.
Add subtitles automatically is the job. preview_caption_style is how you check a BASIC look against your real plate. list_caption_fonts is how you stop guessing at a font id. None of those calls invent a keeper.
Burned in means the wording is final
These subtitles are pixels, not a sidecar you can toggle off. There is no SRT to edit after the fact inside this tool. If the transcript is dirty — names, product SKUs, a legal line — you wanted a plain transcript first (transcribe_audio), a human pass on the text, then a burn. You did not want a full-length subtitle render on audio you already planned to replace.
Transcription quality follows the mix. Heavy noise, cross-talk, and overlapping speakers degrade the cues. Isolating a vocal after you subtitle is backwards: you captioned the noise. Isolate the keeper, or replace the scratch, then subtitle.
BASIC presets are the 1× rate and the right default for accessibility. DYNAMIC presets are 2× and a style choice. If the goal is "people can read this muted," do not spend the premium tier on a candidate.
The first step is the take, not the track
People reach for auto-subtitles because the blank lower third looks unfinished. Unfinished is correct while the shot is still in taste. Generate until the performance and the line survive review. Lock duration. Confirm the spoken language is the language you will ship — this tool transcribes as-is; it does not translate. Then burn.
165 language codes exist. Coverage is not equal. That is another reason not to test language on a throwaway: you will learn the model struggles on this accent, throw the file away, and still have spent the pass.
The AI video generator is where the take comes from. Subtitles are how that take becomes readable. Do not concatenate those jobs in your head. One is sampling. The other is finishing.
A short sequence that holds
- Keep the shot. If native audio is unusable, attach or replace it.
- If names have to be exact, transcribe, correct, then caption. If they do not, a BASIC auto-burn on the keeper is enough.
- Preview a plain style on a few seconds. Render the rest once.
Ask estimate_cost before a long file. A month of clips at DYNAMIC is a different bill from the same month at BASIC, and both are wasted if the soundtrack moves.
FAQ
Are automatic subtitles a different engine from social captions?
No. Same add_veed_captions tool. The intent here is a readable track, so pick a plain preset. The credit rule is identical: do not burn a full pass on a file you will replace.
Can I subtitle while I am still trimming?
You can. You should not. A trim changes in-times. Burned cues that belonged to the discarded head or tail are gone, and cues that remain may sit on the wrong frames. Trim the keeper, then subtitle.
Does this translate into another language?
No. It writes the language that was spoken. A translated soundtrack is audio-only translation or a full dub. Subtitle the result, not the English scratch you are about to throw out.
Why not run it on every generation as a QC check?
Because you will read the cues, hate the take, regenerate, and pay again. Watch the picture. If the take is wrong, the subtitle pass never should have happened.