You write the line. The tool places it. Nobody has to have said it out loud. Captions are the other door: they are a timed transcript of audio. Word-highlight captions still come from speech. Mixing the two is how a CTA you typed gets treated as a subtitle track, or a spoken line gets treated as a poster.
Asking an image model to render the headline inside the still is a third door, and a bad one when the letters have to stay editable. Ideogram can letter a stamp on a product plate; that is still a generate of type, not an overlay you can restyle later.
The editing task is add text overlay to video for one persistent line, or timed text overlays when several lines need their own in/out. Burned-in captions are a different task on the same picture.
In practice
- If nobody said the line, it is overlay, not captions.
- Keep one persistent badge as overlay; put a sequence of claims on timed overlays.
- Do not generate the headline into the still unless the type is part of the product photography.
The mistake to avoid
Sending a mute clip to auto-captions because you wanted a headline. Auto-captions transcribe speech; overlay is copy you wrote.
Where you will run into it
- Add a Text Overlay to a Video — Your words, placed exactly where you want them.
- AI Caption Generator — Drop a video, get timed captions. Free, on your device, no upload.
- AI Reel Maker — Hook, beats, captions, posted.
Related terms
Burned-in captions
Burned-in captions meaning: subtitles rendered into pixels, so viewers cannot turn them off or lose them on re-upload. Closed captions are a separate track.
Word-highlight captions
Word-highlight captions are burned-in subtitles that light up each word as it is spoken, using word-level timing and a caption-style accent colour.
Prompt
Prompt meaning: the written instruction a generative model reads to decide what to make. The one input almost every model needs.
Hook rate
Hook rate meaning: the share of people shown a video who are still watching a few seconds in. It grades the opening, not the body.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync, in your browser or on your phone.
Free on iPhone. On a computer? The same account works on Versely Web, no install needed.