AI voice and music in October 2026: Suno v6
October 2026 audio letter: Suno v6 replaces v5, Gemini 3.8 TTS gets 2,000+ voices, Lyria 3.5 and ElevenMusic v2.5 land. What's on Versely, what's closest.
1 October 2026. Last month's letter was mostly about speech: four TTS rows, Hume's Octave 2 as the outside comparison, and a warning that Suno's licensing story was a music problem filed under voice. Since then the music side has moved more than speech has. Suno retired its own previous models. Google and ElevenLabs both shipped new music models. Google's speech model also jumped a version, and one video lab made its avatars real-time.
Some of the launches below are already on Versely, some are not. Each section says which, and names the closest row you can run today where the answer is no.
Speech: Gemini 3.8 Flash TTS, now the Versely default
Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS on 22 September per the API changelog; press coverage dated it the 23rd. The headline is the voice library, which grows from 30 voices to more than 2,000. It covers 100+ languages, designs a voice from a text prompt, and clones a voice from a 30-second sample. It also handles two-speaker scenes, laughs and sighs, and marks output with SynthID (Gemini API changelog; Unite.AI).
One limit decides a lot of plans: voice cloning is unavailable in the EEA, the UK, Switzerland, India, Illinois and Texas. If your team or your speaker sits in one of those places, the 3.8 clone path is closed to you on Google, whatever you think of the quality.
Both 3.8 models reached the Versely voice studio, web and app, on 24 September, and Flash is now the default voice. What Versely runs is the stock-voice path; Google's voice design and cloning are not wired in. By job:
- Two voices, inline tags, style direction. Gemini 3.8 Flash TTS at 2 credits per 1,000 characters, or Flash Lite at 1 for bulk lines. The 30 stock voices, tags like
[sigh]rendered as sounds, and two-speaker dialogue. The previous default, Gemini 3.1 Flash TTS, is 12. - Clone a real speaker. Cartesia Sonic 3.6 is the newest Cartesia row; last month's letter linked 3.5. The catalog lists 44 languages including Hindi-English code-switching, emotion read from context without tags, speed control and cloning, at 4 credits per 1,000 characters. Voices cloned on earlier Sonic versions carry over.
- High-volume or conversational turns. Inworld TTS 2 Flash is the low-latency tier of Inworld's TTS-2, with cloning, at 2 credits per 1,000 characters.
- Design a voice from a description. Gemini 3.8 Flash TTS does this too: describe how it should sound and it reads the script that way over a stock voice. Versely retired Qwen 3 TTS on 24 September, and voice design moved to 3.8 at the same per-character price.
The 2,000-voice library is the one thing no Versely row matches yet. If you need to audition a big roster before casting, do that on Google, then produce on whichever row holds the voice you picked.
Music: Suno v6 took v5 off Suno's picker
Suno launched v6, v6-wild and v6-mini on 9 September. The new line is trained on licensed catalogues from WMG, BMG and Believe. It edits part of a song by prompt or by lyric, takes text, an image or a video as a reference, and isolates stems. v6-mini is free on Suno; v6 and v6-wild are paid. At the same time v5, v4.5 and v4 came out of Suno's model picker, and users are pushing back (TechCrunch; Digital Music News).
Read the removal for what it is. Suno is moving everyone onto the model trained on licensed music while its lawsuits continue. If your channel's sound was built on a v5 preset, that preset now lives only in the files you already exported.
On Versely, Suno Sounds V6 and its Mini and Wild siblings sit on one page at 2 credits a call. They make sound effects and background music with looping, tempo, key and lyric-subtitle capture. Songs run in the Music studio, where the Suno V6 family is the default. Suno's own song-editing and stem tools are Suno features. Do not expect them on the Sounds rows.
ElevenMusic v2.5 and the UMG deal
ElevenLabs made Music v2.5 its default music model on 11 September, with richer instrumentation. Listeners preferred it in blind tests across 47,885 prompt pairs, API support followed on 14 September, and v2 stays available (Music Business Worldwide).
The Universal Music Group deal signed on 10 September is a separate thing. No UMG music trained v2.5. Our UMG and ElevenLabs explainer covers what the deal announces, a licensed remix platform, and what it does not.
ElevenLabs Music v2.5 is in the Versely Music studio on the web: one prompt, sung or instrumental, up to 10 minutes, with the credit price shown before you generate. For lyrics you wrote yourself, MiniMax Music 3 (next section) is the better fit. For a bed under a cut, Suno Sounds V6.
Lyria 3.5 is GA, and on Versely since 24 September
Google made Lyria 3.5 generally available in the Gemini API and the Gemini app on 3 and 4 September. It writes full songs of about two to three minutes, in 44.1 kHz stereo, as WAV. It takes up to 10 input images and watermarks output with SynthID (Gemini API changelog; Unite.AI on Lyria 3.5). The earlier Flow Music rollout is in our Lyria 3.5 piece.
Lyria 3.5 reached the Versely Music studio, web and app, on 24 September: one prompt, a full song sung or instrumental, at the same credit price as Lyria 3 Pro. Image-to-music stays on Google for now; the Versely row takes a text prompt. The other full-song row, for lyrics you wrote yourself, is MiniMax Music 3, which MiniMax released on 13 August (Releasebot). The Versely row writes complete songs up to five minutes from a style description plus lyrics. Structure tags such as [verse] and [chorus] go on their own lines. It returns 44.1 kHz 16-bit stereo and bills 3 credits per 1,000 characters. There is no image-to-music input on that row, so if the brief is "score this still", that part stays on Google.
The choice between a song and a bed has not changed. A three-minute vocal track is right for a music-led post or a lyric video. For a 20-second product cut, a looped bed from the AI music generator usually fits better, and nobody has to clear lyrics.
Pika Audio (not on Versely)
Pika shipped an audio family on 18 August with four parts: Soundtrack, Music, SFX and Speech. Speech runs at 48 kHz and supports cloning (Pika blog).
None of it is on Versely. For soundtrack and SFX, the closest row is Suno Sounds V6. For cloned speech, Cartesia Sonic 3.6. Music with vocals goes to MiniMax Music 3. On Versely those jobs run as separate rows in one studio.
Lip-sync: Vidu S2-Avatar went real-time (not on Versely)
ShengShu's Vidu S2 landed on 15 September. S2-Avatar is a real-time interactive character at up to 720p, up from 540p, and S2-Editing edits a live stream as it plays (PR Newswire).
Vidu S2 is not on Versely, and none of the Versely lip-sync rows is real-time. The lip-sync and avatar rows here render a file from a face and an audio track:
- Kling Avatar Pro, a talking avatar up to 4K, from 46 credits.
- LTX 2.5 Audio to Video, which times video to a supplied clip of up to 10 seconds: Fast from 52 credits, Pro from 68.
- Avatar X, a script-to-avatar row with the voice included, from 120 credits.
If the product is a live character answering a viewer, you need a real-time platform. If the product is a clip, pick a row with the lip-sync row guide and generate the voice first.
Quiet this month: Hume and Stability
Hume shipped no new model in the window. Octave 2 is still its current speech model, and it is still not on Versely. Stability released no new audio model either. Last month's comparison of Octave 2 against the Versely TTS rows still holds.
October routing
| Job | Versely row | From |
|---|---|---|
| Cloned speaker, many scripts | Cartesia Sonic 3.6 | 4 credits per 1,000 chars |
| Two-hander with tags | Gemini 3.8 Flash TTS | 2 credits per 1,000 chars |
| Cheap high-volume reads | Inworld TTS 2 Flash | 2 credits per 1,000 chars |
| Full song with lyrics | MiniMax Music 3 | 3 credits per 1,000 chars |
| Song from one prompt | Lyria 3.5 or ElevenLabs Music v2.5 (Music studio) | Quoted before you generate |
| Bed or SFX under a cut | Suno Sounds V6 | 2 credits a call |
| Talking avatar from a face | Kling Avatar Pro | 46 credits |
| Video timed to your audio | LTX 2.5 Audio to Video | 52 credits |
Decisions for the month
Export anything you still need from Suno v5. On Suno it is gone from the picker. Your old files are the only copy of that sound.
Check where your speaker lives before planning a Gemini 3.8 clone. Six regions are excluded.
Keep song and bed separate. Full songs go to Lyria 3.5, ElevenLabs Music v2.5 or MiniMax Music 3. Beds go to Suno Sounds V6 or Lyria 3 Clip.
Do not sell real-time. Nothing on Versely does it. Vidu S2-Avatar does.
Plan the month in credits on the rows above, in the web app or on your phone.
FAQ
Can I still use Suno v5?
Not on Suno's own picker, where v6 replaced it on 9 September. The Suno Sounds V5 row you may see in the Versely catalog is the sound-effects line, not the v5 song model. For songs, plan on v6 and keep your exported v5 files.
Is Gemini 3.8 Flash TTS on Versely?
Yes, since 24 September: Flash and Flash Lite are in the voice studio, and Flash is the default. Versely uses the stock voices with style direction, tags and two speakers. For cloning use Cartesia Sonic 3.6; Google's own 3.8 clone feature is unavailable in the EEA, UK, Switzerland, India, Illinois and Texas anyway.
Did UMG's music train ElevenMusic v2.5?
No. The UMG licensing deal is separate from v2.5, and reporting states no UMG music trained it.