The 2026 audio timeline
16 March – 15 April 2026
3 models · 3 providersGoogle, Grok and Suno shipped 3 audio models between 16 March 2026 and 15 April 2026. Of the 3, 2 have a spec page and 1 is folded into a parent model's page as tier or mode variants. Cheapest complete job in the window: 2 credits on Suno Sounds V5.5.
- New names on the roster: Grok, Google.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Grok TTSxAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000… | Grok | 4 credits (headline rate) | — | Text to audio |
| Suno Sounds V5.5variantSuno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle… | Suno | 2 credits | — | Text to audio |
| Gemini 3.1 Flash TTSGoogle Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags… | 12 credits (headline rate) | — | Text to audio |
5 May – 23 June 2026
3 models · 3 providersByteDance, Cartesia and Inworld shipped 3 audio models between 5 May 2026 and 23 June 2026. All 3 have a spec page.
- New name on the roster: ByteDance.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Inworld TTS 2Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language… | Inworld | 2 credits (headline rate) | — | Text to audio |
| Cartesia Sonic 3.5Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and… | Cartesia | 4 credits (headline rate) | — | Text to audio, Voice clone |
| Seed Audio 1.0ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio… | ByteDance | 17 credits (headline rate) | — | Text to audio |
24 August – 2 September 2026
3 models · 3 providersCartesia, Inworld and MiniMax shipped 3 audio models between 24 August 2026 and 2 September 2026. All 3 have a spec page.
- Longest single clip the catalog had offered — 300s on MiniMax Music 3.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| MiniMax Music 3High-performance music generation for complete songs up to five minutes. Takes a style description plus lyrics, with structure… | MiniMax | 3 credits (headline rate) | — | Text to audio |
| Cartesia Sonic 3.6Cartesia's newest and top-ranked TTS model (Sonic 3.6). 44 languages including Hindi-English code-switching, context-driven… | Cartesia | 4 credits (headline rate) | — | Text to audio, Voice clone |
| Inworld TTS 2 FlashInworld Realtime TTS-2 Flash: the low-latency tier of TTS-2, around five times faster to first audio and 40% cheaper, with the… | Inworld | 2 credits (headline rate) | — | Text to audio |
11–23 September 2026
5 models · 2 providersGoogle and Suno shipped 5 audio models between 11 September 2026 and 23 September 2026. Of the 5, 2 have a spec page and 3 are folded into a parent model's page as tier or mode variants. Cheapest complete job in the window: 2 credits on Suno Sounds V6.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Suno Sounds V6Suno Sounds V6 generates sound effects and background music from text prompts with greater detail, plus looping, tempo, key… | Suno | 2 credits | — | Text to audio |
| Suno Sounds V6 MinivariantSuno Sounds V6 Mini is the lightweight, faster cut of V6, balancing quality and speed, with looping, tempo, key controls and… | Suno | 2 credits | — | Text to audio |
| Suno Sounds V6 WildvariantSuno Sounds V6 Wild pushes creative boundaries for bolder, more distinctive sound design, with looping, tempo, key controls and… | Suno | 2 credits | — | Text to audio |
| Gemini 3.8 Flash Lite TTSvariantGoogle Gemini 3.8 Flash-Lite TTS - fast, low-cost speech: 30 voices, 101 languages, style direction, inline audio tags and… | 1 credit (headline rate) | — | Text to audio | |
| Gemini 3.8 Flash TTSGoogle Gemini 3.8 Flash TTS - Google's most expressive voice model: 30 voices, 130 languages, natural-language style direction,… | 2 credits (headline rate) | — | Text to audio |
When 2026 was busy
6 of the twelve months carried a audio release; the busiest window was 11–23 September 2026, with 5.
- March 2026
- 2
- April 2026
- 1
- May 2026
- 1
- June 2026
- 2
- August 2026
- 2
- September 2026
- 6
Who shipped audio in 2026
Suno (4), Google (3) and Cartesia (2) led on volume. SKUs, not quality — four tiers of one model count four times.
| Provider | Models | Spec pages | First | Latest |
|---|---|---|---|---|
| Suno | 4 | 1 | 26 March 2026 | 11 September 2026 |
| 3 | 2 | 15 April 2026 | 23 September 2026 | |
| Cartesia | 2 | 2 | 16 June 2026 | 27 August 2026 |
| Inworld | 2 | 2 | 5 May 2026 | 2 September 2026 |
| ByteDance | 1 | 1 | 23 June 2026 | 23 June 2026 |
| Grok | 1 | 1 | 16 March 2026 | 16 March 2026 |
| MiniMax | 1 | 1 | 24 August 2026 | 24 August 2026 |
What 2026 moved
Firsts, measured against every audio model released before them. “New name on the roster” is the provider label in the catalog, not the company — one lab can hold several labels.
- 16 March – 15 April 2026
- New names on the roster: Grok, Google.
- 5 May – 23 June 2026
- New name on the roster: ByteDance.
- 24 August – 2 September 2026
- Longest single clip the catalog had offered — 300s on MiniMax Music 3.
Dates are each model’s released_at value, walked in order and grouped until a window held 3 or more. Credits are what one complete generation costs per the model’s own price matrix; where it states none, the headline rate is shown and the model sits out the cheapest-in-window line.
Other release years
Frequently asked questions
How many AI audio models were released in 2026?+
Versely's catalog carries 14 audio models with a 2026 release date, from 7 providers, arriving in 4 launch windows between 16 March 2026 and 23 September 2026.
What was the biggest AI audio launch of 2026?+
11–23 September 2026, with 5 audio models from Google and Suno. Of the 5, 2 have a spec page and 3 are folded into a parent model's page as tier or mode variants. Cheapest complete job in the window: 2 credits on Suno Sounds V6.
Which company released the most AI audio models in 2026?+
Suno, with 4 of the 14 audio models dated 2026 — first on 26 March 2026, most recently on 11 September 2026. Google shipped 3, Cartesia shipped 2, Inworld shipped 2.
What changed in AI audio generation in 2026?+
Measured against everything the catalog carried before it: New names on the roster: Grok, Google; New name on the roster: ByteDance; Longest single clip the catalog had offered — 300s on MiniMax Music 3.
Run any 2026 audio model in Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.