Eleven v4 and Eleven v4 Turbo are the latest text to speech models from ElevenLabs. Eleven v4 is described as the company's most emotive model yet, built on an entirely new architecture that reads a script the way a voice actor would, understanding who is speaking, what just happened, and how every line should land. Eleven v4 Turbo is the fastest real-time speech model with the expressive range of Eleven v4, designed for voice agents and live conversations. Both models are available in the ElevenLabs apps, through the Eleven API, and through ElevenAgents.
The core problem these models address is the difficulty of creating synthetic speech that sounds genuinely human and emotionally expressive rather than flat or robotic. Traditional text to speech often struggles with consistent speaker identity, natural pacing across long scripts, and conveying emotion or context. Eleven v4 tackles these issues by using a new architecture that interprets the script like an actor, following direction tags for emotions and sound effects, maintaining speaker stability across regenerations, and using context stitching to keep pacing steady across long-form content like audiobooks. For real-time applications, latency has been a major barrier; Eleven v4 Turbo directly addresses this with a median inference latency of approximately 100 milliseconds and a median time to first speech of approximately 150 milliseconds.
Eleven v4 includes several key features. Write direction into the script functionality allows users to add tags like [laughs], [whispers], and [door slams] directly into the text, and the model follows these tag sequences more reliably than the previous v3 model, including sound effects. Context stitching keeps pacing and delivery consistent across any script length, so a full audiobook sounds like a single take from the first page to the last. Regenerate without vocal drift ensures that redoing a line once or fifty times still results in the same person speaking, with speaker stability holding across dialogue, narration, and everything in between. Professional Voice Clones, which were not supported in v3, are back in Eleven v4 and perform with the model's full emotional range across every language they speak.
Eleven v4 Turbo is optimized for live conversation and real-time use cases. Stream in, stream out functionality lets text be pushed as an LLM generates it, with audio starting to come back before the sentence is finished, using bidirectional streaming built for agent loops. The model carries the full expressive range of Eleven v4, so confirmations, escalations, and holds land differently from one another instead of reading identically. Professional Voice Clones work identically across both v4 and v4 Turbo, keeping a single brand voice consistent from the first turn of a call to the last. The Turbo model also handles multilingual speech fluently, such as Japanese, Spanish, or Portuguese, with a native accent.
Both models also benefit from the broader ElevenLabs voice ecosystem. Users can cast from over 17,500 voices in the voice library, organized into categories such as Narration, Conversational, Social media, Character, Educational, Advertisement, Entertainment, and Multilingual voices. Each voice category is designed for a specific purpose, from warm and authoritative narration voices to energetic short-form social media voices and trustworthy educational voices. This gives creators a wide range of choices to match the tone and context of their content.
The products include voice customization options. Voice Cloning allows cloning a voice from ten seconds of audio, or training a Professional Voice Clone for a near-perfect match. Voice Design lets users describe a voice in a sentence and generate it without needing a recording session or casting call. The Pronunciation Dictionary allows defining how names, acronyms, and technical terms are pronounced, with phonetic settings applied to every generation, such as OAuth as oh-auth, Hülkenberg as hool-ken-berg, and Reykjavik as rayk-yah-veek.
Eleven v4 outputs to the same audio formats as every ElevenLabs model, including MP3 for podcasts, YouTube, and general listening; WAV or PCM for studio work, dubbing, and post-production; and µ-law optimized for telephony and call-center integrations. Sample rate and bitrate are set via the API, allowing audio to be tuned for quality or bandwidth depending on the destination. A single generation supports up to 10,000 characters, with context stitching supporting longer content.
The benefits for users include highly expressive and controllable synthetic speech, reduced latency for real-time interactions, consistent speaker identity across long projects, and support for 90+ languages. The models are available on every plan including a free tier with 10,000 credits per month (roughly 10 minutes of audio), and premium plans starting at $6 per month that include 30,000+ credits, professional voice cloning, and higher limits. Enterprise plans offer custom pricing, higher volume, custom SSO, and priority support.
Use cases explicitly described in the content include audiobook narration that sounds like a single take, real-time voice agents for customer-facing calls, voice experiences in games through companies like Rosebud, publishing journalism to audio with higher engagement and longer listening times, conversational dialogue and podcasts that feel unscripted, short-form social media content for platforms like TikTok and Reels, characters for games and animations, educational courses and training simulations, and advertisement reads for digital ads or TV and radio.
The target users include podcasters, audiobook creators, game developers, educators, marketers, advertisers, social media creators, publishers, and developers building voice agents or real-time applications. The API supports REST endpoints, streaming endpoints, and TypeScript and Python SDKs. The models are trusted by enterprise customers and are SOC 2 Type II certified, ISO 27001 certified, PCI DSS Level 1 certified, GDPR compliant, and support HIPAA-eligible workflows for healthcare.
Eleven v4 and Eleven v4 Turbo represent a significant step forward in text to speech technology, combining high emotional expressiveness with low-latency performance for real-time use. Whether the goal is producing polished long-form content or powering live voice agents, these models aim to deliver natural, controllable, and consistent speech across languages and use cases.