It's one half of a two-context problem: the same viewer watches muted on a train and with earbuds in at home, and a video ships into both without knowing which one it'll land in. Burned-in captions solve the muted half by making sure nothing is lost with the sound off. Sound-on strategy is the other half — making sure something is actually gained when a viewer does have it on, rather than the video being identical either way.
Native-audio generation makes this more achievable than it used to be, because dialogue, effects and ambience come out of the same generation as the picture rather than being bolted on afterward in a separate pass — a beat can land on the audio because both were generated to agree with each other from the start.
The two halves aren't in tension so much as sequenced: build the muted version first since it has to work for the largest share of a feed audience, then layer the sound-on reward on top, rather than writing a voiceover-dependent script and hoping captions can rescue it later.
In practice
- Design the muted version first — captions, on-screen text, visual pacing — then add what sound-on viewers specifically gain on top of it.
- Use trending or native audio where a platform's algorithm favours it, rather than an original score alone, when discovery matters more than a distinctive score.
- Test with sound off before shipping, even for content meant to be watched with sound on — most of the audience will still see it muted first.
Models that generate their own audio
Catalog entries flagged as producing sound alongside picture. 113 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Wan 3.0 Text to Video | Wan | Video |
| Seedance 2.0 | ByteDance | Video |
| Grok Imagine Video | Grok | Video |
| Pixverse 5.6 Text to Video | Pixverse | Video |
| Kling Video V3 Pro Text to Video | Kling | Video |
| Vidu Q3 Image to Video | Vidu | Video |
| Vidu Q3 Video | Vidu | Video |
| Pixverse 5.6 Image to Video | Pixverse | Video |
Browse all 56 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Writing a voiceover-carried script and treating captions as the fix for muted viewing. A caption track transcribes what was said; it doesn't recreate the pacing a visual-first cut would have had without the voiceover propping it up.
Go deeper
Designing Brand Videos for Sound-On and Sound-Off Viewing
How to design brand videos that work muted and rewarded with sound on: dual-track structure, platform sound defaults, and native-audio AI models.
Related terms
Burned-in captions
Burned-in captions meaning: subtitles rendered into pixels, so viewers cannot turn them off or lose them on re-upload. Closed captions are a separate track.
Native audio
Native audio meaning: a video model that generates soundtrack (dialogue, effects, ambience) in the same pass as the picture.
Hook rate
Hook rate meaning: the share of people shown a video who are still watching a few seconds in. It grades the opening, not the body.
Spark Ads
Spark Ads meaning: TikTok budget behind an existing post (yours or a creator's) so paid reach keeps the real likes, comments, and shares.
Whitelisting
Whitelisting meaning: a brand running paid ads through a creator handle (with permission) so the ad shows their name, not the brand's.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.