Speech, voice and audio

    Audio tags

    Also called Emotion tags, Inline delivery cues.

    Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.

    They exist because delivery is positional. A single "sound excited" setting applies to a whole take, but real reads change within a sentence: the aside is quieter, the punchline is faster, the sigh lands between two clauses. A tag sits at the exact point the change should happen.

    The vocabulary is provider-specific and so is the syntax. Some use square brackets, some angle brackets, some expect a paired opening and closing tag around the phrase the emotion applies to. There is no shared standard, so a script marked up for one model is not portable to another.

    Crucially, not every model supports them, and on a model that does not, tags are simply text. It will read the brackets aloud. Models without inline support usually take direction through a separate style or emotion field instead, which applies to the whole take.

    In practice

    • Place a tag immediately before the phrase it modifies, not at the top of the script.
    • Paired tags must be closed, or the delivery change runs to the end of the take.
    • Use them sparingly — a tag on every clause produces a read that lurches.

    The mistake to avoid

    Copying a tagged script between providers. On a model that does not parse them the brackets are pronounced, and on one with a different vocabulary they are ignored.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.