Speech, voice and audio

    Voice stability

    Also called Stability, Exaggeration, Expressiveness.

    Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.

    It governs variance, not quality. A high-stability setting keeps every take close to a neutral, even read, which is what you want for long-form narration where consistency across twenty paragraphs matters more than any single line landing hard. A low setting lets the model swing further: more dynamic range, more emotional colour, and more chance that one line comes out oddly.

    Providers name the same idea differently — stability on one, exaggeration on another, expressiveness elsewhere — and some invert the direction, so the same numeric value can mean opposite things across two models. Read what the control is called and which way it runs before carrying a number across.

    It interacts with everything else you use to direct the read. Inline tags and emotion settings have less room to operate at high stability, because you have just told the model to stay near the middle.

    In practice

    • High stability for narration and anything long; lower for character work and short punchy lines.
    • Change one control at a time — stability and emotion settings mask each other.
    • Values do not port between providers, even when the label matches.

    The mistake to avoid

    Turning expressiveness up to fix a boring read. Past the middle of the range you mostly get inconsistency, not conviction.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.