On-screen text in generated video: what holds up
Text a video model draws inside the frame has to be right 200 times, not once. The size threshold where letterforms break, and the clean-plate rule.
Two completely different things get called "text on the video." One is a title card, caption or lower third that a timeline draws on top of finished footage. The other is text the video model itself renders inside the scene: a shop sign, a product label, a t-shirt, a phone screen, a book cover. The first is a solved problem. The second is one of the last places generated video still reliably falls over.
If you have read about legible text in still images, none of it transfers cleanly. A still has to be right once. A clip has to be right every frame, and that changes the maths so completely that a model with excellent still typography can still produce unusable video signage.
The 200-frame problem
Versely renders at 25 frames per second by default rather than 24, so an 8-second clip is 200 frames. Every one is a fresh opportunity for a letterform to change.
Assume, generously, that a model renders your word correctly in 99 out of every 100 frames. Treating the frames as independent, which is a simplification but a directionally honest one, the chance the whole clip is clean is 0.99 to the power of 200, or about 0.13. Roughly seven clips in eight will contain at least one bad frame.
That is the entire reason still-image typography benchmarks mislead here. A model can be genuinely excellent at glyphs and still fail this, because the requirement is not accuracy, it is accuracy sustained across two hundred consecutive independent draws. This is a temporal consistency problem wearing a typography costume.
And a bad frame in text is not like a bad frame elsewhere. A softened hand for two frames is invisible. A word that reads BAKERY for 180 frames and BAKREY for 6 is the thing the viewer's eye locks onto, because reading is involuntary.
The size threshold, and why it is about stroke width
The useful threshold is not font size, it is stroke width in output pixels.
A letterform is a structure of strokes. If the strokes are thick enough to occupy several pixels, there is real structure for the model to carry between frames. If a stroke is two pixels wide, it is inside the range the model re-derives from noise on every frame, and it will flip.
Work backwards. For a regular-weight sans-serif, cap height is roughly seven times stroke width. A four-pixel stroke, which is about the floor at which there is real structure to carry, implies a 28-pixel cap height. In a 1080-pixel-tall frame that is about 2.6 percent of frame height. In a 720p frame the identical 2.6 percent is only a 19-pixel cap height and a stroke under three pixels, so the same proportion has already dropped below the floor.
That gives you a working rule, stated against a 1080p frame:
- Under about 3 percent of frame height: at or under the stroke floor, so expect drift. This is most real signage in a real environment.
- 3 to 8 percent: unstable. Sometimes fine at one seed, wrong at the next.
- Above about 8 percent: the model has enough structure to hold, and this is where the occasional convincing generated sign comes from.
Notice what that rules out. A shop sign at readable distance, a nutrition panel, a phone UI, a business card, small print on packaging — all of them live below the threshold. The generated text that works in demo reels is almost always enormous, flat-on, and on screen briefly.
Three things that make it worse, in order
Motion of the lettered surface. Static text on a static surface is one problem. Text on a surface that rotates, tilts or passes under a moving camera has to be re-projected under changing perspective every frame, and the model is solving typography and geometry simultaneously. A logo on a turning bottle is the standard worst case, and it is worst case by a wide margin.
Length. Every additional word is an additional set of glyphs that can independently go wrong, and the failure rate compounds. One short word above the size threshold has a real chance. A full sentence of body copy does not, at any size.
Specificity. This is the one people underrate. If the word is decorative and nobody knows what it should say, a wrong letter costs nothing. If it is your brand name, your product name or a price, one wrong letter is not a soft failure. It is a client escalation. Text you cannot afford to get wrong should never be generated inside the frame, regardless of what the size arithmetic says.
What actually holds
Given the above, a short list of in-frame text that is worth attempting:
| Case | Why it survives |
|---|---|
| A single large word, flat to camera, under 3 seconds | Above the stroke threshold, few frames, no re-projection |
| Deliberately illegible background signage | No correct answer exists, so no wrong answer is visible |
| Non-Latin decorative text where legibility is not the point | Same reason |
| A number on a jersey or door, large and static | Short glyph set, high stroke width |
And the counterpart, which is longer: brand names, product labels, prices, URLs, phone numbers, book titles, UI screens, subtitles, anything on a moving surface, anything more than two or three words.
For still images the calculus is different and better, which is why the standard workflow is to generate a still with correct typography on a model built for it, then animate that still. Seedream 5 Pro carries typography, multilingual text and dense-layout capability in its catalog entry, and Ideogram v4 is the other obvious candidate. Animating a still that already has correct text is not immune to drift, but it starts from a correct frame one instead of inventing letterforms from scratch, and it is a materially better bet than pure text-to-video.
The clean-plate rule
The rule that resolves nearly all of this: generate the plate empty, stamp the text in the timeline.
Prompt for a blank sign, an unlabelled bottle, a bare wall, a phone with a dark screen. Then add the words as a layer in the editor. What you get:
- Pixel-identical letterforms in every frame. The layer is drawn once and composited, so there is no drift to have.
- Editable without regenerating. Change the price, fix the typo, swap the language. The generation is untouched.
- Exact brand type. Overlay text is set from a named font registry rather than from a model's impression of what your typeface looks like.
- Timing control. Words can appear and leave on cue with timed text overlays, which is not something an in-frame render can do at all.
The editor is EDL-based, so the plate and the text layer live on one re-renderable timeline. The 480p preview pass is free and carries a short per-user cooldown, which makes it cheap to check placement and timing before committing. The final export is charged once regardless of how many clips are in the timeline, and there are no watermarks on any plan, so the delivered frame is the frame you designed.
Two practical additions. First, put text, letters, watermark and signage in the negative prompt field where the model exposes one, because unrequested hallucinated text in a background is a common way a plate stops being clean. Second, place overlay text with platform chrome in mind — safe zones covers where a caption survives the interface.
A worked comparison. Same brief, two routes: a 6-second shot of a coffee bag on a counter, brand name visible.
Route A, generated in frame. The prompt names the brand. The bag reads correctly at frame 1, loses a letter around frame 70, regains it, and renders a different letterform for the last second. The brand line sits at maybe 4 percent of frame height, inside the unstable band. Three seeds produce three spellings. Unusable, and no phrasing fixes it.
Route B, clean plate. The prompt asks for an unbranded kraft coffee bag on a counter, slow push-in, with text and logos excluded via the negative prompt field. The plate holds. The brand mark goes on as a timeline overlay in the actual brand font, checked at 480p before export. Every frame identical.
Route B is not a workaround. It is how this shot is done in conventional production too, where the label is a print job and the bag is the camera department's problem.
FAQ
Are there video models that render text well?
Some are visibly better than others, and it is worth testing your specific word. But "better" here means the unstable band moves, not that it disappears. No current video model gives you the frame-to-frame stability a compositing layer has by construction, and betting a client deliverable on in-frame brand text is a risk with no upside, because the timeline route is faster anyway.
Can I fix a few bad frames instead of regenerating?
Only if the text is static and flat. Then you can patch the region across the affected frames. If the surface is moving, you are tracking and replacing per frame, which is more work than generating a clean plate and stamping the word once.
Does this apply to captions and subtitles?
No, and that is the good news. Captions are drawn by the timeline, not by the model, so none of the 200-frame arithmetic applies. They are pixel-stable by construction. See burned-in captions for how they differ from sidecar files.
Is this different from text in still images?
Yes, substantially. Stills have to be right once, so the model classes and prompting advice are genuinely different — that is covered separately in legible text in AI images and in text baked into the image versus stamped over it.