Guides

    The opening text card as a second hook

    On-screen text in frame one hooks independently of the audio and can quietly contradict the visual. How to test card variants against one identical cut.

    Versely Team9 min read

    Most creators write one hook and ship three. The spoken line is the one they agonise over. The opening image is the one they choose. And the text card burned across frame one — the six words sitting over the top third — usually gets typed in at the end of the edit, in about four seconds.

    That card is doing hook work whether or not it was designed to. A large share of the audience reads it before hearing anything, because a large share of the audience has the sound off.

    Frame one is three channels, and they can disagree

    The evidence on sound-off viewing is unusually solid for this field. XR (Extreme Reach) and MX8 Labs surveyed more than 3,000 people across the US, UK, France, Spain and Germany in 2026. Among US respondents, 49% said they use captions always or often, a behaviour the study described as consistent across screens and genres. Coverage of the same study put always-or-often caption use on short-form at 39%. Earlier work by Verizon Media with Publicis, an April 2019 survey of 5,616 US adults aged 18 to 54, found that 80% of respondents said they were more likely to watch a video through when captions were available, and that 50% agreed captions matter because they usually watch with the sound off.

    So frame one carries three independent hooks:

    • The visual. What the first frame shows.
    • The audio. The first spoken clause, heard by the subset with sound on.
    • The text. The card, read by everyone, first, and fastest.

    Reading is faster than listening. A viewer who lands on your clip has taken in six words of overlay before your first sentence has finished its second syllable. On the sound-off path, the card is the hook and the audio is a bonus track.

    The contradiction nobody catches

    The failure mode is specific and it is invisible in the edit suite, because in the edit suite you have watched the clip forty times with sound on and you already know what it is about.

    Three shapes recur:

    The card labels, the visual promises. The frame shows something mid-collapse; the card says "Kitchen renovation part 3". The visual took on a debt and the text paid it off as an admin note.

    The card promises, the visual is inert. The card says "This ruined the whole build"; the frame is a person sitting in a chair about to talk. The text has written a cheque the image cannot cash within a second.

    The card and the audio say the same thing twice. The voice says "here's why your thumbnails aren't working" and the card says "why your thumbnails aren't working". This is redundancy, and it wastes the only channel that could have added a second fact.

    The third one is the most common and the cheapest to fix. Two channels, two facts. If the audio states the topic, the card should state the stake, the number, or the surprise.

    None of this is captioning. Captions transcribe what was said; an overlay carries the claim. They are different tools with different jobs, and conflating them is what produces card three above. If that distinction is fuzzy, the subtitles versus overlays breakdown is the cleaner treatment.

    Writing a card that hooks

    Constraints first, because they are not negotiable:

    Constraint Value Why
    Line length 42 characters or fewer Broadcast subtitling convention, and it survives a phone screen
    Lines on screen 2 maximum Three lines reads as a paragraph and gets skipped
    Hold Long enough to read twice A card that leaves before a re-read is a card nobody finished
    Vertical placement Clear of the bottom ~15% and top ~10% Platform UI sits there and will cover you

    Within that box, the card has one job: add a fact the other two channels do not carry. Some patterns that hold up:

    • The number the voice hasn't said yet. Visual shows the workshop, voice says "we tried something stupid", card says "£40 of materials".
    • The stake. Card names what is at risk, voice narrates the attempt.
    • The negative frame. "What to avoid" beats "what to do" in most creative testing, and a card is the cheapest place to test that framing because it changes nothing else about the video.
    • The unresolved half. Card poses the question, the video answers it. This only works if the answer arrives soon; a card that opens a loop the body takes ninety seconds to close is the same overselling problem in a different channel.

    Skip the ones that read as metadata: episode numbers, series names, your own handle. Those belong later in the clip, or in the caption field where they cost nothing.

    Typography and contrast decide whether any of this is legible over moving footage, and that is a separate craft problem covered in the typography and hierarchy guide and the readability rules. Pick a face from a font registry rather than eyeballing it — choosing from the registry explains why guessing produces a different look on every upload.

    The test: one cut, several cards

    The reason text cards are the best variable on the board is that they are the only hook channel you can change with absolute confidence that nothing else moved.

    Change a spoken hook and you have changed the audio, which changes the waveform, which changes the caption timings, which sometimes changes the first cut point. Change the opening frame and you have changed the visual. Change the card and you have changed six words of pixels.

    On an EDL-based timeline this is structural rather than aspirational. The timeline is a set of instructions, so the video, the cuts, the audio and the caption track stay byte-identical across variants while the overlay text differs. Render variant A, swap the card, render variant B. Watch each through the free 480p preview pass first — it carries a short per-user cooldown, so use it to check placement and legibility rather than to browse — and the final export is charged once per render regardless of how many clips the timeline contains.

    Then run it properly:

    1. Build three cards, not two. Two variants tests a wording tweak. Three lets you test categories — a number card, a stake card, a negative-frame card — which is where real differences live.
    2. Never post variants at the same time. They compete for the same audience pool and cannibalise each other. Same posting slot, consecutive days, is the standard.
    3. Name the metric before you look. For an opening card the honest metric is the swipe-away rate or average watch time at a fixed 24-hour checkpoint. Not views. Views on day one are mostly feed lottery.
    4. Pre-commit a stop rule. Early data skews toward whichever variant posted first and reverses with volume. Decide the checkpoint, then look once.
    5. Run at least three pairs before concluding anything. A single pair of short-form posts is dominated by distribution variance, not by your copy.

    The full discipline, including sample-size honesty, is in the creative A/B protocol.

    Reading the result

    Two outcomes are informative and one is a trap.

    A card wins on the swipe metric and holds through the body. Adopt the category, not the sentence. If the number card won, write number cards.

    A card wins the swipe and loses at ten seconds. The card overpromised relative to the video. Repair the body or the claim, not the card design.

    The variants land within noise of each other. This is the trap, because the temptation is to declare the winner anyway. Within-noise means the card is not currently the binding constraint on this clip. Move to a different variable — the opening frame, or the first spoken clause — rather than iterating on wording that is not moving anything.

    One more thing worth logging as you go: whichever card family wins, keep the losers. A tested card library is reusable across a whole content pillar. Auto-generated caption tracks handle the transcription layer separately, and the caption generator will not touch your overlay.

    FAQ

    Does the text card matter on platforms where most people watch with sound on?

    Sound-on and sound-off are not clean segments — the same person switches by context. XR described caption use as consistent across screens rather than clustering on mobile, so treating the card as decoration means the sound-off portion of your audience gets a hook you never wrote. The card is cheap enough that there is no case for leaving it to chance.

    How long should the card stay up?

    Long enough for a comfortable second read, and off before it starts competing with the next beat. Broadcast subtitling conventions put cue holds in the one-to-six-second range, which is a reasonable outer bound, but an opening card is closer to the short end because the video needs to move.

    Can I test the card and the opening frame in the same experiment?

    You can run them as separate arms, but not as one changed variant. If both differ between A and B, a result tells you nothing about which channel caused it. Sequence them: card first because it is cheapest, frame second.

    What if the card is the only thing that changed and the numbers still swing wildly?

    That is the expected behaviour of a single short-form pair, and it is why the three-pair rule exists. Distribution variance on any one post is larger than most copy effects. If three consecutive pairs point the same direction, you have something; if they alternate, you have noise wearing a result's clothes.