Gemini 3.8 TTS: voice design and cloning
Gemini 3.8 Flash TTS adds voice design, 30-second cloning, two-speaker scenes and a far bigger voice library. Who it suits, where cloning is blocked.
Google put two new speech models into the Gemini API on 22 September 2026: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts (Gemini API changelog). Press coverage dated it a day later (Unite.ai). The headline is not "better voices". It is that a Google TTS model now does voice design, voice cloning and staged two-person scenes, jobs that used to need three separate vendors.
One thing up front. Both 3.8 models are in the Versely voice studio, on the web and in the app, since 24 September: Gemini 3.8 Flash TTS is the new default voice. What Versely runs is the stock-voice path with style direction, audio tags and two speakers. Google's voice design and cloning are not part of it yet, so the last section maps those two jobs to the rows that do them today.
What shipped
Two models, one split. Flash TTS is the flagship. The changelog describes "studio-grade voice fidelity, nuanced acting, regional dialects". Flash-Lite TTS is the "fast, cost-efficient" tier for high-throughput production. Pick Lite when you generate thousands of short lines. Pick Flash when a person will listen to the whole read.
Languages. More than 100 languages and dialects, including regional varieties such as Mexican Spanish, Quebec French and Scots English (Unite.ai). Regional accents matter more than the language count. A Mexican Spanish ad read in a Castilian accent is a wrong read.
The voice library. This is where the sources disagree. Unite.ai reports the library going from 30 voices to "more than 2,000 production-ready voices". Google's own changelog summary says "150+ voices in extended library". Both agree it is a big step up from 30. Until you can browse the list in AI Studio, plan around "many more voices", not an exact count.
Voice design. Describe a voice in plain language (role, accent, character) and the model builds it. You no longer have to scroll a voice list hoping one fits.
Voice cloning. It takes a 30-second sample of your own voice. Unite.ai reports built-in consent verification, SynthID watermarking and C2PA credentials on cloned output.
Two-speaker scenes. Native two-speaker staging directs a multi-turn conversation in one request. You can direct each line separately, and script vocal bursts such as <laughs>, <sigh> and <gasp>.
Long-form stability. Google claims voice quality, pacing and timbre hold steady "across hours". That is a vendor claim. Test it on your own long script before you rely on it.
Watermark. Output carries SynthID, which Google describes as imperceptible and embedded in the audio.
Scores. Unite.ai reports Flash TTS scored 71.4 on Hume AI's Voice Design Benchmark and 60.8 on accent modelling. There is no independent leaderboard run yet. Treat those as launch numbers.
Where cloning is blocked
Cloning is not available everywhere. Per Unite.ai, it is unavailable in:
| Region | Cloning |
|---|---|
| European Economic Area | Blocked |
| United Kingdom | Blocked |
| Switzerland | Blocked |
| India | Blocked |
| Illinois (US) | Blocked |
| Texas (US) | Blocked |
Voice design and the stock library are not on that list. Only cloning is. Illinois and Texas both have biometric-privacy laws, and the EEA and UK have GDPR, so the list reads like a legal map, not a technical one. That is our reading, not Google's stated reason.
If your studio or your client sits in one of those places, do not build a cloning workflow on 3.8 and hope a VPN sorts it out. That is how a client contract breaks. Our rule on voice-clone consent as a client clause applies either way.
Price and availability
Both models went live the same day in the Gemini API and Google AI Studio. Flash TTS also appears in Gemini Notebook and Flash-Lite TTS in Google Vids. Enterprise access is "coming soon" (Unite.ai).
Google's pricing page lists a launch rate that ends on 31 December 2026:
| Model | Text input | Audio output | From 1 January 2027 |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.50 | $9.00 | $1.00 in / $18.00 out |
| Gemini 3.8 Flash-Lite TTS | $0.50 | $6.00 | $1.00 in / $12.00 out |
Those are Google's API list prices, per 1M tokens on the standard tier. Output prices double in January. If you are pricing a long audiobook or a daily narration series, cost it at the 2027 rate, not the launch rate.
Who it is for
Developers shipping voice inside a product. Voice design, cloning and two-speaker scenes in one API means one integration instead of three. This group gains the most.
Localisation teams. Regional dialects are the real upgrade. If you ship the same ad into Mexico, Quebec and Scotland, the accent control is worth testing first.
Long-form narrators and podcasters. Hours-long stability plus two-speaker staging is aimed straight at audiobooks and explainer podcasts. Test a 30-minute chapter before you commit a season.
Not ideal for: a creator in the EEA, UK, Switzerland, India, Illinois or Texas whose plan rests on cloning. It is also a poor fit for anyone who needs a finished video rather than an audio file. TTS gives you a voice track. A talking picture is still a separate lip-sync or native-audio step.
Our recommendation. If you build in the API and your users sit outside the blocked regions, prototype on Flash-Lite, move hero reads to Flash, and budget at the 2027 price. If you make videos rather than apps, you do not need to switch vendors this week. The same three jobs already run elsewhere.
Gemini 3.8 on Versely, and the jobs it does not cover yet
What is live. Gemini 3.8 Flash TTS (2 credits per 1,000 characters on the catalog card) and Gemini 3.8 Flash Lite TTS (1 credit) are in the voice studio's model picker. Flash is the default. Both take a natural-language style direction, render inline tags such as [sigh] and [laughs] as sounds rather than words, and stage two-speaker dialogue. They read from the 30 stock voices. Google's bigger library, its voice design and its cloning are API features Versely has not wired in. Versely's own voice design runs on 3.8's style direction, and cloning runs on the rows below. Credits are read from each model row.
Voice design: Gemini 3.8 Flash TTS. Describe the voice in words ("gravelly older narrator, slow and confident") and it reads your text that way, over one of the 30 stock voices. The same job is a sentence in the agent: design a custom AI voice. Since 24 September this runs on 3.8 and bills per character, like any other read; the Qwen 3 TTS rows that used to do it are retired. One limit to know: it does not create a reusable voice ID. Keep the description and the stock voice name in your brief so the sound stays the same from script to script.
Cloning: Cartesia Voice Clone (8 credits) or Inworld Voice Clone (2 credits). Cartesia is the path the agent uses: attach a recording, name it, and clone my voice from a recording returns a reusable voice you call by name from then on. Clean the sample first. A noisy room gives you a noisy clone. Clone only voices you have the rights to, wherever you are.
Two-speaker scenes: Gemini 3.8 Flash TTS (2 credits per 1,000 characters). It does multi-speaker dialogue, natural-language style control and inline audio tags, and it costs a sixth of the 3.1 row it replaced as the default (Gemini 3.1 Flash TTS is 12). It works from 30 preset voices. For a scripted podcast intro or an interviewer-and-guest exchange, write each line with a speaker alias and ask the agent to create a multi-voice dialogue. You get one audio file back, in speaking order.
When the stock voices are enough. Short ads, UGC voiceovers, two-person skits and anything under a minute. The preset voices and tags cover that ground. When you will feel the gap: a regional accent that no preset matches, or a named cast larger than 30 distinct voices. For a bigger roster today, Inworld TTS 2 (2 credits) is the style-steerable option with broad language coverage.
What to do this week
- If you build on the Gemini API: try Flash-Lite on your highest-volume line, Flash on your hero read, and cost both at the January 2027 rate.
- If your users are in a blocked region: keep cloning off the 3.8 plan entirely. Use design or stock voices there.
- If you make videos: move your reads to Gemini 3.8 Flash TTS in the voice studio, Lite for bulk lines. Design and stage dialogue on 3.8, clone on Cartesia. For the wider field, state of AI voice, September 2026 covers the other engines.
The 3.8 launch matters because it pulls three jobs under one vendor. It does not change which job you are doing. Pick the job, then the row.