Text Overlays: Typography and Hierarchy on Video
Text overlay rules for video that actually get read: hierarchy, safe zones, timing math, contrast fixes, and overlay vs caption decisions for short-form.
Watch any high-performing Reel with the sound off — which is how a huge share of viewers watch it anyway — and count the words on screen in the first two seconds. Almost always between three and seven. Not zero, not a paragraph. The text overlay is doing the hook's job while the video does the mood's job, and the accounts that get this split right routinely out-retain accounts with objectively better footage.
Text on video is typography under hostile conditions: a moving, unpredictable background, a viewer deciding in under a second, and a canvas partially eaten by platform UI. Most overlay advice ("use bold fonts!") ignores the actual failure modes, which are hierarchy, placement, timing, and contrast — in that order.
Here is the system I use across every short-form video, whether the footage is AI-generated or shot on a phone.
Hierarchy: one job per text element
Every piece of text on a frame has exactly one of three jobs, and each job has its own size class:
- Hook text — the reason to keep watching. Biggest element on screen, 3 to 7 words, on frame within the first second. "I stopped buying stock footage" is a hook. "Welcome to my channel" is not.
- Context labels — the small supporting facts: a price, a step number, a product name, "day 3 of 30." Roughly half the hook's size. One at a time.
- Captions — the spoken-word transcript riding along the bottom third. Smallest class, consistent position, never competing with the other two.
The most common overlay failure is three elements at the same size fighting for the same eyeball. If everything is loud, nothing is. Pick the hook, make it unmissable, and demote everything else visibly.
Captions are their own discipline with their own tooling — auto-timing, styled presets, word-level emphasis — covered in the auto caption generator guide. This post is about the two deliberate classes above them.
Safe zones: where platforms eat your text
Platform UI covers more of the frame than most people internalize. On a 9:16 vertical:
| Zone | What lives there | Verdict for text |
|---|---|---|
| Top ~10% | Search bar, "Following/For You" tabs | Avoid |
| Right edge, lower half | Like/comment/share rail | Avoid — costs you ~20% of width |
| Bottom ~15% | Caption text, username, audio attribution | Avoid for overlays; captions sit just above it |
| Center 60% band | Nothing | Prime real estate |
Practical rule: keep hook text in the upper-middle of the center band, context labels near the subject they describe, captions just above the bottom UI zone. Design at 9:16 and check the frame with a platform UI mockup once — after you have seen your hook half-hidden behind a share button one time, you never skip the check again.
Timing: the 1.5-words-per-second rule
Viewers read on-screen text at roughly 3 words per second when relaxed — and they are not relaxed, they are mid-scroll. Budget conservatively:
- Hold time = word count ÷ 1.5, minimum 1.5 seconds. A five-word hook needs about 3 seconds on screen.
- One text change per beat. If your video cuts every 2 seconds, text can change every 2 seconds. Text changing faster than the edit feels like spam; slower feels stale.
- Never make text and motion compete. Big camera move plus new text at the same moment means one of them goes unread. Land text on the settled half of a shot.
This math is why paragraph-length overlays fail: a 25-word overlay needs 16+ seconds of hold, which no short-form shot sustains. If you need 25 words, that is five sequential overlays or a voiceover.
Contrast: making text survive a moving background
A static design can tune text color to its background once. Video backgrounds move, so your text needs armor that works everywhere:
- Scrim or solid tag. A semi-transparent dark box behind white text is the most reliable option and the visual language of native TikTok text. Ugly-effective beats pretty-invisible.
- Outline plus shadow. White fill, thin dark outline, soft shadow — survives most backgrounds without a box. This is the standard caption-preset construction.
- Frame for the text at generation time. The AI-native option: if you are generating footage anyway, prompt for negative space ("subject on the right third, clean sky behind on the left") so the text has a calm home. You control the background, which flat-footage editors never could.
One thing to stop doing: baking text into the image generation prompt for overlays. Even models with strong typography — Seedream 5 Pro renders text in 14 languages and is genuinely good at it — produce baked-in text you cannot retime, resize, or A/B test. Generated typography is for posters and thumbnails; video overlay text should stay a separate, editable layer.
Style: two fonts, two colors, every video
Overlay styling collapses to a short set of rules that cover 95% of brand content:
- Two fonts maximum — one heavy sans for hooks, one regular weight for labels and captions. Same pair in every video; the consistency itself becomes brand recognition.
- Two colors plus white. A brand accent for emphasis words, white for everything else, and the scrim. Rainbow text reads as 2019.
- Sentence case beats ALL CAPS for anything over four words. Caps hooks work at 3 or 4 words; caps sentences slow reading measurably.
- Animate entrances, not existence. A quick pop or slide-in draws the eye; text that pulses continuously exhausts it.
In Versely, the text overlay tool positions styled text on any generated or uploaded video, and the same hierarchy applies in the AI slideshow maker, where each slide is effectively a held frame with a text budget — slideshow text tolerates about double the word count because the frame holds still.
Where overlays fit in the layer stack
Order of operations for a finished short: footage first, then picture-in-picture video layers, then text overlays, then captions. Text goes on after any video overlay layers so it stays readable above them, and captions go last so their timing maps to final audio. Retiming text because you added a layer underneath it afterward is a self-inflicted wound — I know because I inflict it on myself about once a month.
FAQ
How many words should a text overlay have?
Three to seven for a hook, and hold it on screen for word count divided by 1.5 seconds, with a 1.5-second floor. Anything longer than about ten words should be split into sequential overlays or moved into the voiceover.
Where should I place text on a vertical video?
In the center 60% band, ideally upper-middle for hooks. The top 10% is covered by platform navigation, the right rail by engagement buttons, and the bottom 15% by captions and attribution. Check your frame against a platform UI overlay once per template.
What's the difference between text overlays and captions?
Captions transcribe spoken audio and ride in a consistent bottom-third position with auto-timing. Overlays are deliberate editorial text — hooks and labels — that you place and time by hand. They coexist in one video but must occupy different size classes and zones.
Should I generate text inside the AI video or add it after?
Add it after, as a separate layer. Baked-in generated text cannot be retimed, edited, or localized, and a typo means a full regeneration. Reserve generated typography for static assets like posters and thumbnails where the text is part of the artwork.
What makes text readable over changing video backgrounds?
A semi-transparent scrim box is the most reliable; outline-plus-shadow is the cleaner-looking second choice. Best of all, prompt your AI footage with deliberate negative space so the text lands on a calm region — an option unique to generated video.
Put the system to work: generate footage with room for the message in the AI video generator, layer your hooks and labels with Versely's text overlays, and let styled auto-captions carry the transcript — free credits daily.