The alternative is closed captions: a separate track the player draws, which the viewer can disable and which every platform styles its own way. Short-form video went to burn-in almost universally because both of those properties are wrong for a muted autoplay feed — you cannot afford the caption to be off, and the caption's look is part of the edit rather than a preference.
The render is three steps that people collapse into one word. Speech is transcribed, the transcript is aligned to timings, and the styled text is composited onto the frames. Errors at each step look different: a wrong word is transcription, a caption that arrives late is alignment, and text sitting under the platform's own interface is layout.
One word also covers two different jobs, and Versely keeps them apart. Transcribed subtitles follow whatever is being said and are generated automatically from the audio, with a set of styling presets and a font list to choose from. A fixed overlay — a hook line, a headline, a badge, a call to action — is text you wrote, placed for a stretch of the clip, and transcribes nothing.
In practice
- Keep captions clear of the platform's interface: the bottom strip of a vertical frame is covered by buttons and account names.
- Read the transcript before rendering — product names and brand spellings are the words transcription gets wrong.
- Pick one caption style per channel; the look is as recognisable as the voice reading it.
The mistake to avoid
Burning in captions before the cut is locked. They are pixels once rendered, so a later trim means recaptioning the whole clip rather than nudging a track.
Where you will run into it
- Add Captions to a Video — Speech in, styled subtitles out — no manual timing.
- Add Subtitles Automatically — You upload the video. The words appear, timed to the voice.
- Add a Text Overlay to a Video — Your words, placed exactly where you want them.
- AI Caption Generator — Transcribed, styled, and rendered into the picture.
- AI Reel Maker — Hook, beats, captions, posted.
Related terms
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Forced alignment
Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.
Hook rate
Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.
Retention curve
A retention curve shows how many of the people who started a video are still watching at each second of it, and watch time is the area under that curve.
UGC ad
A UGC ad is an advert made to look like an ordinary person's own post — handheld, spoken to camera, unpolished on purpose — so it reads as a recommendation rather than a commercial.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.