Generation modes

    Text overlay

    Also called Title overlay, On-screen text, Hook text.

    Text overlay is copy you authored, burned onto a clip at a position you chose, rather than a transcript of speech or a new generate that painted letters into the scene.

    You write the line. The tool places it. Nobody has to have said it out loud. Captions are the other door: they are a timed transcript of audio. Word-highlight captions still come from speech. Mixing the two is how a CTA you typed gets treated as a subtitle track, or a spoken line gets treated as a poster.

    Asking an image model to render the headline inside the still is a third door, and a bad one when the letters have to stay editable. Ideogram can letter a stamp on a product plate; that is still a generate of type, not an overlay you can restyle later.

    The editing task is add text overlay to video for one persistent line, or timed text overlays when several lines need their own in/out. Burned-in captions are a different task on the same picture.

    In practice

    • If nobody said the line, it is overlay, not captions.
    • Keep one persistent badge as overlay; put a sequence of claims on timed overlays.
    • Do not generate the headline into the still unless the type is part of the product photography.

    The mistake to avoid

    Sending a mute clip to auto-captions because you wanted a headline. Auto-captions transcribe speech; overlay is copy you wrote.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.