They exist because delivery is positional. A single "sound excited" setting applies to a whole take, but real reads change within a sentence: the aside is quieter, the punchline is faster, the sigh lands between two clauses. A tag sits at the exact point the change should happen.
The vocabulary is provider-specific and so is the syntax. Some use square brackets, some angle brackets, some expect a paired opening and closing tag around the phrase the emotion applies to. There is no shared standard, so a script marked up for one model is not portable to another.
Crucially, not every model supports them, and on a model that does not, tags are simply text. It will read the brackets aloud. Models without inline support usually take direction through a separate style or emotion field instead, which applies to the whole take.
In practice
- Place a tag immediately before the phrase it modifies, not at the top of the script.
- Paired tags must be closed, or the delivery change runs to the end of the take.
- Use them sparingly — a tag on every clause produces a read that lurches.
The mistake to avoid
Copying a tagged script between providers. On a model that does not parse them the brackets are pronounced, and on one with a different vocabulary they are ignored.
Where you will run into it
- Add a Voiceover to a Video — Type the script. Get a narrated video back.
Related terms
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Voice stability
Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.