AI Models

    Eleven v4 and v4 Turbo: release, price, vs v3

    ElevenLabs released Eleven v4 and v4 Turbo on 28 September 2026: 90+ languages, audio tags, Turbo at ~150 ms, API list price $0.08 per 1,000 characters.

    Versely Team••8 min read

    ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026. Both cover 90+ languages and follow inline audio tags. v4 is the quality model; Turbo is the real-time one, with about 100 ms median inference latency. The API lists v4 at $0.08 per 1,000 characters, with a launch discount that ran to 12 October. Neither is in Versely's catalog.

    Key facts

    Item Detail
    Released 28 September 2026
    Variants Eleven v4 (eleven_v4) for content and long-form audio. Eleven v4 Turbo (eleven_v4_turbo) for real-time use
    Where ElevenCreative (creator apps), ElevenAgents (voice agents) and ElevenAPI. A free account can try both. OpenRouter also lists elevenlabs/eleven-v4-turbo
    API price, listed 8 October 2026 v4: $0.08 per 1,000 characters. Turbo: $0.04. Launch discount of 72% ran to 12 October
    App plans Free tier with 10,000 credits a month. Paid plans start at $6 a month; commercial use needs one. ElevenLabs gives no separate v4 credit rate
    Speed Turbo: ~100 ms median inference latency, excluding app and network time; ~150 ms median time to first speech in ElevenLabs' own test
    Limits v4: 10,000 characters per request, about 10 minutes. No SSML, no Style or Speed sliders
    Cloning Instant clones from as little as 10 seconds of audio, ElevenLabs says. Professional clones supported

    Sources, all read on 8 October 2026: ElevenLabs' launch post, model docs, changelog and API pricing; Artificial Analysis' Provider Voice and Controlled Voice boards; and TechTimes (30 September 2026), which corroborates the launch figures. The performance numbers are theirs, not Versely's own tests.

    What changed from Eleven v3 and Multilingual v2

    ElevenLabs' docs call v4 "a net upgrade over Eleven v3, delivering better results in almost every case". The documented lineup, with Elo from Artificial Analysis' Provider Voice board on 8 October 2026:

    Model Languages Chars per request Elo
    v4 Turbo 90+ 10,000* 1332
    v4 90+ 10,000 1324
    v3 70+ 5,000 1174
    Multilingual v2 29 10,000 1097
    Flash v2.5 32 40,000 1079

    * Not in ElevenLabs' docs table; OpenRouter lists 10,000 for v4 Turbo.

    Tags and control. ElevenLabs says v4 follows audio tags like [laughs], [said angrily in French accent] and [light rain] "more accurately than prior models". Its docs are franker: tags are "not perfect yet", and a tag can occasionally be read as a sound effect instead of a delivery cue, so write [low, gravelly voice] rather than something that sounds like a noise. Pauses come from ellipses and punctuation, because v3 and v4 don't support SSML break tags. IPA pronunciation support is improved.

    Languages. The docs list 91 languages for v4, against 74 for v3 and 29 for Multilingual v2. Additions since v3's list include Cantonese, Mongolian, Odia, Amharic and Zulu. In ElevenAgents, ElevenLabs says Turbo's biggest gains are in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

    Cloning and accents. ElevenLabs says an Instant Voice Clone can capture a voice "with high fidelity using just 10 seconds of audio". Professional Voice Clones now work with v4; you train an existing one on v4 from My Voices. Its docs describe a typical Instant sample as one to two minutes and say source quality matters more than before. Cross-language output also changed on purpose: generate in a language other than the reference voice's and v4 aims for fluent, native-sounding speech instead of carrying the source accent over. A v4 clone "may sound quite different from its Eleven v3 version".

    Dialogue. Multi-speaker scenes go through the Text to Dialogue endpoints, and the docs say speakers are uncapped. The launch changelog added continuity options (preceding and following text, request IDs), and ElevenLabs calls request stitching "significantly more reliable".

    Checking the headline claims

    #1 on Artificial Analysis. True for native voices. On the Provider Voice board on 8 October 2026, Turbo scores 1332 and v4 1324, both with a rank range of 1 to 2, ahead of Qwen-Audio-3.1-TTS-Plus (1295), Cartesia Sonic 3.6 (1279) and Gemini 3.8 Flash TTS (1276). (TechTimes had v4 at 1319 on 30 September.) The Controlled Voice board, which clones the same eight English voices for every model, is less clear-cut: Qwen-Audio-3.1-TTS-Plus is first at 1184, with Turbo (1169) and v4 (1159) close behind.

    About 75% of listeners prefer it. This is ElevenLabs' own blind test against Cartesia Sonic 3.6, Inworld TTS-2 and two Gemini 3.8 models. Graders heard the same line from v4 and one rival, with ties counted as half. Its chart says v4 wins 65% to 81% of its blind head-to-head tests. No sample size is published, and TechTimes called it a "company-run evaluation".

    About 150 ms to first speech. Median for Turbo over WebSocket, measured by ElevenLabs in September 2026 with network latency removed. Its chart puts the rivals it tested, among them Cartesia Sonic 3.6 and OpenAI GPT-4o mini TTS, at 262 to 814 ms. The ~100 ms in the docs is inference only. Neither figure includes your network or the rest of an agent's pipeline.

    v4 or v4 Turbo?

    Pick v4 for anything you render offline: narration, audiobooks, dubbing, game dialogue. It runs through Create speech, Stream speech and the Text to Dialogue endpoints, takes 10,000 characters a request, and ElevenLabs calls it its highest-quality model.

    Pick Turbo for live conversation. The docs route eleven_v4_turbo through the Text to Dialogue WebSocket, and the 5 October agents API changelog lists it as the new default conversational voice model, replacing eleven_flash_v2. It lists at half v4's price.

    Don't expect a large quality gap: on both Artificial Analysis boards the two scores sit inside each other's margins of error.

    Who it is for, and who should wait

    Good fit: creators who direct delivery line by line (audiobooks, character reads, game dialogue), teams dubbing one voice into many languages, and voice-agent builders who want expressive speech at low latency. v3 and Multilingual v2 list at the same $0.08 per 1,000 characters, so v4 is not a price rise.

    Test first, or wait, if:

    • You rely on SSML or the Style and Speed sliders. v4 has neither; Stability and Similarity are its only voice settings.
    • A cloned voice must keep its native accent in other languages. That is now the default, and a toggle is still "a research project".
    • You have tuned v3 voices. Clones can sound different, and ElevenLabs says v4's behavior "may shift over time" as it keeps training.
    • Price per character decides it. v4 lists at $80 per million characters; Artificial Analysis lists Cartesia Sonic 3.6 at $49 and Gemini 3.8 Flash TTS at $16.50 for the same million.

    Can you use it on Versely?

    No. As of 8 October 2026 neither Eleven v4 nor Eleven v4 Turbo is in Versely's model catalog, and we can't say whether or when that will change. Versely's ElevenLabs speech rows are Multilingual v2 at 8 credits per 1,000 characters and a 4-credit single-language Speech Turbo row. Neither is v4.

    For the same job today:

    Job Use Credits Elo
    Tagged reads, two-speaker dialogue Gemini 3.8 Flash TTS (default) 2 1276
    Style direction, 100+ languages Inworld TTS 2 2 1252
    Cloned voices, 44 languages Cartesia Sonic 3.6 4 1279
    ElevenLabs sound, 29 languages ElevenLabs Multilingual 8 1097

    Credits are per 1,000 characters. Elo is Artificial Analysis' Provider Voice board on 8 October 2026. Inworld TTS 2 is listed there as Realtime TTS-2.

    Three of those voices were in ElevenLabs' own preference test. In Artificial Analysis' blind votes, v4 and Turbo sit about 45 to 80 Elo points above the first three rows. Cloning on Versely is free to create and billed per 1,000 characters when the clone speaks, through Cartesia or Inworld rather than ElevenLabs: see AI voice cloning. For the ElevenLabs row against the default voice, read ElevenLabs vs Gemini Flash TTS.

    FAQ

    When was Eleven v4 released?

    ElevenLabs released both models on 28 September 2026. The launch post carries that date and was last updated on 5 October.

    How much does Eleven v4 cost?

    On the API, ElevenLabs lists v4 at $0.08 per 1,000 characters and Turbo at $0.04, as read on 8 October 2026. A 72% launch discount ($0.022 and $0.011) ran to 12 October 2026. App plans start free, with 10,000 credits a month, and paid plans from $6 a month.

    What are the Eleven v4 API model IDs?

    eleven_v4 for Create speech, Stream speech, Create dialogue and Stream dialogue, and eleven_v4_turbo for the Text to Dialogue WebSocket.

    Is Eleven v4 better than Eleven v3?

    ElevenLabs says so, and Artificial Analysis' blind votes agree: v4 scores 1324 against 1174 for v3 on Provider Voice, and 1159 against 1071 on Controlled Voice. It also doubles the request limit to 10,000 characters at the same list price. Clones may sound different from their v3 versions.

    Can I use Eleven v4 on Versely?

    No. As of 8 October 2026 it is not in Versely's catalog. Use Gemini 3.8 Flash TTS (2 credits per 1,000 characters), Inworld TTS 2 (2), Cartesia Sonic 3.6 (4) or ElevenLabs Multilingual (8).