Most caption families drop a whole line as a block, timed to a phrase. Two families — Snap and Reveal — are timed: each spoken token takes the accent colour, and on Snap a pop scale, at the moment alignment says that word started. The rest of the line stays dim until its turn. Variants inside those families change the accent, not the clock.
Burned-in captions is the burn itself — pixels in the file, not a sidecar track. Forced-alignment is the clock. Speech-to-text is the transcript. Word-highlight is the look those two timed families actually vary, which is what the caption-styles catalog is selling when you pick Snap Gold versus Reveal Cyan.
Apply them on the transcribe-and-burn path. A silent product plate has no word to time. An authored overlay that nobody said is not karaoke — that is text you wrote, on a different tool.
In practice
- Choose Snap for a fast hook or ad read; choose Reveal for a slower, suspenseful voiceover.
- Judge the family by its base variant, then pick an accent; colour is not a different caption system.
- Keep timed words clear of the mouth and the platform's bottom chrome — Snap lives at the bottom and will hit both.
The mistake to avoid
Restyling a subtitle family until it 'feels like karaoke,' or burning Snap onto a silent clip. Kind is timed or it is not, and a plate with no speech has nothing to highlight.
Go deeper
Nine caption families across 45 presets
Versely's 45 caption presets are nine families of five. Map them by backing, position and word timing to shortlist three candidates in a couple of minutes.
Where you will run into it
- Add Captions to a Video — Speech in, styled subtitles out — no manual timing.
- AI Caption Generator — Transcribed, styled, and rendered into the picture.
Related terms
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
Forced alignment
Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Hook rate
Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.
Retention curve
A retention curve shows how many of the people who started a video are still watching at each second of it, and watch time is the area under that curve.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.