It governs variance, not quality. A high-stability setting keeps every take close to a neutral, even read, which is what you want for long-form narration where consistency across twenty paragraphs matters more than any single line landing hard. A low setting lets the model swing further: more dynamic range, more emotional colour, and more chance that one line comes out oddly.
Providers name the same idea differently — stability on one, exaggeration on another, expressiveness elsewhere — and some invert the direction, so the same numeric value can mean opposite things across two models. Read what the control is called and which way it runs before carrying a number across.
It interacts with everything else you use to direct the read. Inline tags and emotion settings have less room to operate at high stability, because you have just told the model to stay near the middle.
In practice
- High stability for narration and anything long; lower for character work and short punchy lines.
- Change one control at a time — stability and emotion settings mask each other.
- Values do not port between providers, even when the label matches.
The mistake to avoid
Turning expressiveness up to fix a boring read. Past the middle of the range you mostly get inconsistency, not conviction.
Related terms
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.