2026 releases · AI audio

    Every AI audio model released in 2026

    10 models from 7 providers, dated 21 January 2026 to 23 June 2026. They did not arrive evenly — they landed in 3 bursts, and this page is those bursts in order.

    The 2026 audio timeline

    21–22 January 2026

    4 models · 2 providers

    Inworld and Qwen shipped 4 audio models between 21 January 2026 and 22 January 2026. Of the 4, 3 have a spec page and 1 is folded into a parent model's page as tier or mode variants.

    • New name on the roster: Qwen.
    ModelBuilt byCredits per jobMax outputCapabilities
    Inworld TTS 1.5 MaxInworld Realtime TTS 1.5 Max — the #1 ranked Inworld model, delivering the best balance of quality and speed. Expressive,…Inworld4 credits (headline rate)Text to audio
    Qwen 3 TTS 0.6BCompact text-to-speech model with natural voice synthesis and efficient processingQwen2 credits (headline rate)Text to audio
    Qwen 3 TTS 1.7BvariantHigh-quality text-to-speech model with enhanced naturalness and emotional expressionQwen3 credits (headline rate)Text to audio
    Qwen 3 TTS Voice DesignQwen 3 text-to-speech with custom voice designQwen3 credits (headline rate)Text to audio, Voice clone

    16 March – 15 April 2026

    3 models · 3 providers

    Google, Grok and Suno shipped 3 audio models between 16 March 2026 and 15 April 2026. All 3 have a spec page. Cheapest complete job in the window: 2 credits on Suno Sounds V5.5.

    • New names on the roster: Grok, Google.
    ModelBuilt byCredits per jobMax outputCapabilities
    Grok TTSxAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000…Grok4 credits (headline rate)Text to audio
    Suno Sounds V5.5Suno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle…Suno2 creditsText to audio
    Gemini 3.1 Flash TTSGoogle Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags…Google4 credits (headline rate)Text to audio

    5 May – 23 June 2026

    3 models · 3 providers

    ByteDance, Cartesia and Inworld shipped 3 audio models between 5 May 2026 and 23 June 2026. All 3 have a spec page.

    • New name on the roster: ByteDance.
    ModelBuilt byCredits per jobMax outputCapabilities
    Inworld TTS 2Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language…Inworld5 credits (headline rate)Text to audio
    Cartesia Sonic 3.5Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and…Cartesia5 credits (headline rate)Text to audio, Voice clone
    Seed Audio 1.0ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio…ByteDance4 credits (headline rate)Text to audio

    When 2026 was busy

    5 of the twelve months carried a audio release; the busiest window was 21–22 January 2026, with 4.

    January 2026
    4
    March 2026
    2
    April 2026
    1
    May 2026
    1
    June 2026
    2

    Who shipped audio in 2026

    Qwen (3), Inworld (2) and ByteDance (1) led on volume. SKUs, not quality — four tiers of one model count four times.

    ProviderModelsSpec pagesFirstLatest
    Qwen3222 January 202622 January 2026
    Inworld2221 January 20265 May 2026
    ByteDance1123 June 202623 June 2026
    Cartesia1116 June 202616 June 2026
    Google1115 April 202615 April 2026
    Grok1116 March 202616 March 2026
    Suno1126 March 202626 March 2026

    What 2026 moved

    Firsts, measured against every audio model released before them. “New name on the roster” is the provider label in the catalog, not the company — one lab can hold several labels.

    Dates are each model’s released_at value, walked in order and grouped until a window held 3 or more. Credits are what one complete generation costs per the model’s own price matrix; where it states none, the headline rate is shown and the model sits out the cheapest-in-window line.

    Other release years

    Frequently asked questions

    How many AI audio models were released in 2026?+

    Versely's catalog carries 10 audio models with a 2026 release date, from 7 providers, arriving in 3 launch windows between 21 January 2026 and 23 June 2026.

    What was the biggest AI audio launch of 2026?+

    21–22 January 2026, with 4 audio models from Inworld and Qwen. Of the 4, 3 have a spec page and 1 is folded into a parent model's page as tier or mode variants.

    Which company released the most AI audio models in 2026?+

    Qwen, with 3 of the 10 audio models dated 2026 — first on 22 January 2026, most recently on 22 January 2026. Inworld shipped 2, ByteDance shipped 1, Cartesia shipped 1.

    What changed in AI audio generation in 2026?+

    Measured against everything the catalog carried before it: New name on the roster: Qwen; New names on the roster: Grok, Google; New name on the roster: ByteDance.

    Run any 2026 audio model in Versely

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.