Guides

    Karaoke captions vs static blocks, by format

    Word-by-word burn-in holds in short-form and fatigues in long-form. The per-format rule, and how to catch caption-style cost at specific timestamps.

    Versely Team9 min read

    Word-by-word burn-in is the right caption for a 20-second talking-head Short and the wrong caption for a 12-minute tutorial. The same style that gives a muted scroller something to track becomes a strobe once the viewer has settled in to watch. Two-line static blocks have the opposite shape: they look timid in a feed and they are what long-form actually holds. The format picks the style. If you have been running one preset across both, the retention graph already knows.

    This is not an argument about whether to caption. XR Extreme Reach's 2026 five-country study (US, UK, France, Spain, Germany) found always-or-often caption use at 49% in the US, with similar rates in Spain and France. The Measure's write-up of that study puts always-or-often use on short-form at 39%. An earlier Verizon Media / Publicis survey of 5,616 US adults (2019) found that 80% of consumers said they were more likely to watch a video to the end when captions were available, and that 50% of consumers said captions matter because they watch with the sound off. The open question is which drawing of the words you burn in, and whether that drawing is costing you at particular timestamps.

    Three styles, not two

    "Karaoke" gets used for two different drawings. They do not belong in the same bucket.

    Style What is on screen Reads as Default home
    Word-by-word pop One word at a time, often scaled or recoloured as it is spoken A feed hook. High tracking, high fatigue. Short-form under ~45 seconds, ads, punchy UGC
    Karaoke highlight A full line visible, active word coloured A lyric video. Tracking without the strobe. Short talking-head, education clips that stay under a couple of minutes
    Two-line static block Up to two lines, held, then replaced A subtitle. Readable, almost invisible as a style. Long-form, interviews, 16:9, anything past a few minutes

    All three, once committed, are burned-in captions: pixels in the file, not a player toggle. That is why the style decision is expensive to reverse and cheap to preview.

    On Versely, spoken-word burn-in goes through add_veed_captions. The 21 BASIC presets are line-level: the whole cue arrives together. The 9 DYNAMIC presets (glass, whisper, glide, fusion and the rest) are word-level animated, billed at 2× the standard credit rate. That split is the live version of the table above, not a recommendation to run a DYNAMIC look on a 20-minute video. Basic versus dynamic caption presets is the billing and treatment map.

    Broadcast subtitling conventions that survived into social, and that static blocks should still obey: at most 42 characters per line, at most two lines on screen, a hold of one to six seconds per cue. Word-by-word pop ignores the hold rule on purpose. That is fine for twelve seconds. It is not fine for twelve minutes.

    Keep the drawing out of the platform furniture: bottom ~15% and top ~10% of a 9:16 frame are where interface chrome sits. A lower-third that looks correct in the editor and vanishes under the username is not a style problem, it is a safe-zone problem.

    The per-format rule

    Match the drawing to the job the viewer is doing, not to the tool you opened.

    Short-form, feed, sound often off. Word-timed captions rescue the muted viewer and give the eye a moving target. That is why word-by-word burn-in is the short-form default. If the clip stretches past a minute, switch from word-by-word pop to karaoke highlight (full line, active word coloured). Pop is a hook device. Hooks that last a minute and a half are just noisy.

    Long-form, click-in, sound often on. The viewer already chose the video. Captions are now a transcript and an accessibility layer, not a tracking toy. Two-line static blocks, high contrast, held long enough to read, are the drawing that does not compete with the picture. Word-by-word pop on a talking-head essay reads as a second, faster video laid over the one they clicked. Karaoke highlight is the compromise for long-form education that still wants a little tracking; it is not the default.

    The awkward middle: a three-minute Short, a 90-second Reel, a YouTube video that is really a Short. Length is the discriminator, not the upload slot. If average view duration on the format is under a minute, word-timed. If it is several minutes, blocks.

    Captions are not overlays. Captions transcribe. Overlays carry the claim, the number, the CTA. Mixing those jobs, a karaoke caption that is also trying to be a title card, produces the worst of both: unreadable as a transcript, weak as a poster. Put the claim on a separate text layer and let the caption track stay a caption track.

    How to tell if the style is costing you at a timestamp

    You do not need a native caption A/B. You need the retention graph and a list of where the drawing gets densest.

    1. Mark the caption storms. Play the video with your eyes on the captions, not the picture. Write down every stretch where word-by-word pop is flashing faster than you can comfortably read, or where a static cue is sitting over a shot change, or where a line wraps mid-name. Those timestamps are suspects. Fast speech is the usual producer of storms.
    2. Read the curve at those marks, not at the averages. Open the retention graph and look at the ten seconds around each suspect. A drop that sits on a caption storm and on a tangent, an ad read, or a format break is probably the content. A drop that sits on a caption storm and on a sentence that is otherwise doing its job is the drawing. The discriminator is whether the same kind of sentence, later, with calmer captions, holds.
    3. Check the first three seconds separately. On short-form, word-timed captions in the open are often the thing that stops the swipe. A drop at 0:02 with no captions is a different diagnosis from a drop at 1:40 with a strobe. Do not "fix" a working hook because the body is tiring.
    4. Preview the alternative on the offending stretch, not on a demo reel. Five seconds of your actual footage, two styles. A DYNAMIC look such as glass and a BASIC look such as simple, sampled on the darkest and brightest frames in that window. Previewing a caption style on five seconds of your own video exists specifically so you do not burn a full render to find out the preset disappears over a white wall. Judge legibility first, taste second.
    5. Restyle the window, not the channel. Long-form can carry blocks for the body and a word-timed open, if the first five seconds are doing feed work in a cut-down. Short-form can carry pop for the hook and a highlight style once the promise has landed. A single preset across a 12-minute file is the habit that creates the 1:40 cliff.

    Two related failures get blamed on style and are not style:

    • Reading speed. A correct transcript that dumps fourteen words on screen for half a second is a chunking failure. Caption readability is that layer: line breaks, contrast, hold time. Switching from karaoke to blocks will not save a cue that is up for 400 milliseconds.
    • Timing drift. Word-timed styles expose timestamp error. A block that holds for two seconds hides a 300-millisecond offset that karaoke makes obvious. If glass looks "late" and simple does not, you may have a timing problem rather than a tier problem. Fix alignment before you abandon the family.

    The production path for getting captions onto the file is adding captions. Burn-in is the point of no cheap return, which is why the five-second sample belongs before the full export.

    FAQ

    Can I run word-by-word captions on a long YouTube video?

    You can. Viewers who last past a couple of minutes will feel it as noise. Two-line static blocks are the long-form default for a reason: the viewer is reading a transcript, not being held by a moving target. If you want tracking on education content, use a full-line karaoke highlight, not one-word pop, and still switch to blocks once the runtime is measured in minutes.

    Do captions always lift watch time?

    They lift it for the people who need them: muted viewers, noisy rooms, accessibility, names and numbers the ear misses. XR's 2026 study puts US always-or-often caption use at 49%, and The Measure reports 39% on short-form specifically. The Verizon Media / Publicis finding is that 80% of surveyed consumers said they were more likely to finish when captions were available. None of that is a licence to pick the loudest drawing. Style can spend the lift the captions themselves earned. The longer watch-time case is in why captions hold attention.

    How do I check a specific timestamp?

    Mark the caption storm, then look at retention in the ten seconds around it. If the content at that moment is clean and the drop is aligned with the flashing, restyle that window and re-export. Preview the alternative on those five seconds of footage first, not on a stock demo.

    Are DYNAMIC presets the same thing as static blocks with animation?

    No. They change the pace of reading. DYNAMIC looks highlight or reveal individual words in sync with the audio. BASIC presets update a whole line at a time. Pick on delivery speed and format length, not on which accent colour you like.