You write the line. The tool places it. Nobody has to have said it out loud. Captions are the other door: they are a timed transcript of audio. Word-highlight captions still come from speech. Mixing the two is how a CTA you typed gets treated as a subtitle track, or a spoken line gets treated as a poster.
Asking an image model to render the headline inside the still is a third door, and a bad one when the letters have to stay editable. Ideogram can letter a stamp on a product plate; that is still a generate of type, not an overlay you can restyle later.
The editing task is add text overlay to video for one persistent line, or timed text overlays when several lines need their own in/out. Burned-in captions are a different task on the same picture.
In practice
- If nobody said the line, it is overlay, not captions.
- Keep one persistent badge as overlay; put a sequence of claims on timed overlays.
- Do not generate the headline into the still unless the type is part of the product photography.
The mistake to avoid
Sending a mute clip to auto-captions because you wanted a headline. Auto-captions transcribe speech; overlay is copy you wrote.
Where you will run into it
- Add a Text Overlay to a Video — Your words, placed exactly where you want them.
- AI Caption Generator — Transcribed, styled, and rendered into the picture.
Related terms
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
Word-highlight captions
Word-highlight captions are burned-in subtitles that light up each word as it is spoken, using word-level timing and a caption-style accent colour.
Prompt
A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.