The 2026 audio timeline
21–22 January 2026
4 models · 2 providersInworld and Qwen shipped 4 audio models between 21 January 2026 and 22 January 2026. Of the 4, 3 have a spec page and 1 is folded into a parent model's page as tier or mode variants.
- New name on the roster: Qwen.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Inworld TTS 1.5 MaxInworld Realtime TTS 1.5 Max — the #1 ranked Inworld model, delivering the best balance of quality and speed. Expressive,… | Inworld | 4 credits (headline rate) | — | Text to audio |
| Qwen 3 TTS 0.6BCompact text-to-speech model with natural voice synthesis and efficient processing | Qwen | 2 credits (headline rate) | — | Text to audio |
| Qwen 3 TTS 1.7BvariantHigh-quality text-to-speech model with enhanced naturalness and emotional expression | Qwen | 3 credits (headline rate) | — | Text to audio |
| Qwen 3 TTS Voice DesignQwen 3 text-to-speech with custom voice design | Qwen | 3 credits (headline rate) | — | Text to audio, Voice clone |
16 March – 15 April 2026
3 models · 3 providersGoogle, Grok and Suno shipped 3 audio models between 16 March 2026 and 15 April 2026. All 3 have a spec page. Cheapest complete job in the window: 2 credits on Suno Sounds V5.5.
- New names on the roster: Grok, Google.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Grok TTSxAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000… | Grok | 4 credits (headline rate) | — | Text to audio |
| Suno Sounds V5.5Suno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle… | Suno | 2 credits | — | Text to audio |
| Gemini 3.1 Flash TTSGoogle Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags… | 4 credits (headline rate) | — | Text to audio |
5 May – 23 June 2026
3 models · 3 providersByteDance, Cartesia and Inworld shipped 3 audio models between 5 May 2026 and 23 June 2026. All 3 have a spec page.
- New name on the roster: ByteDance.
| Model | Built by | Credits per job | Max output | Capabilities |
|---|---|---|---|---|
| Inworld TTS 2Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language… | Inworld | 5 credits (headline rate) | — | Text to audio |
| Cartesia Sonic 3.5Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and… | Cartesia | 5 credits (headline rate) | — | Text to audio, Voice clone |
| Seed Audio 1.0ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio… | ByteDance | 4 credits (headline rate) | — | Text to audio |
When 2026 was busy
5 of the twelve months carried a audio release; the busiest window was 21–22 January 2026, with 4.
- January 2026
- 4
- March 2026
- 2
- April 2026
- 1
- May 2026
- 1
- June 2026
- 2
Who shipped audio in 2026
Qwen (3), Inworld (2) and ByteDance (1) led on volume. SKUs, not quality — four tiers of one model count four times.
| Provider | Models | Spec pages | First | Latest |
|---|---|---|---|---|
| Qwen | 3 | 2 | 22 January 2026 | 22 January 2026 |
| Inworld | 2 | 2 | 21 January 2026 | 5 May 2026 |
| ByteDance | 1 | 1 | 23 June 2026 | 23 June 2026 |
| Cartesia | 1 | 1 | 16 June 2026 | 16 June 2026 |
| 1 | 1 | 15 April 2026 | 15 April 2026 | |
| Grok | 1 | 1 | 16 March 2026 | 16 March 2026 |
| Suno | 1 | 1 | 26 March 2026 | 26 March 2026 |
What 2026 moved
Firsts, measured against every audio model released before them. “New name on the roster” is the provider label in the catalog, not the company — one lab can hold several labels.
- 21–22 January 2026
- New name on the roster: Qwen.
- 16 March – 15 April 2026
- New names on the roster: Grok, Google.
- 5 May – 23 June 2026
- New name on the roster: ByteDance.
Dates are each model’s released_at value, walked in order and grouped until a window held 3 or more. Credits are what one complete generation costs per the model’s own price matrix; where it states none, the headline rate is shown and the model sits out the cheapest-in-window line.
Other release years
Frequently asked questions
How many AI audio models were released in 2026?+
Versely's catalog carries 10 audio models with a 2026 release date, from 7 providers, arriving in 3 launch windows between 21 January 2026 and 23 June 2026.
What was the biggest AI audio launch of 2026?+
21–22 January 2026, with 4 audio models from Inworld and Qwen. Of the 4, 3 have a spec page and 1 is folded into a parent model's page as tier or mode variants.
Which company released the most AI audio models in 2026?+
Qwen, with 3 of the 10 audio models dated 2026 — first on 22 January 2026, most recently on 22 January 2026. Inworld shipped 2, ByteDance shipped 1, Cartesia shipped 1.
What changed in AI audio generation in 2026?+
Measured against everything the catalog carried before it: New name on the roster: Qwen; New names on the roster: Grok, Google; New name on the roster: ByteDance.
Run any 2026 audio model in Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.