5 min read

Yesterday's Top Launches: 1 Tools from October 4, 2026

ElevenLabs launched Eleven v4 and Eleven v4 Turbo, two speech models designed to make text-to-speech sound more like a voice actor interpreting a script.

Yesterday's Top Launches: 1 Tools from October 4, 2026

Daily Digest: October 5, 2026

Yesterday was a thin day on the release calendar, with exactly one launch worth writing about. ElevenLabs pushed out two speech models at once, Eleven v4 and Eleven v4 Turbo, and given how much of the current wave of new developer tools is built on voice interfaces, the timing is hard to ignore.

Eleven v4 and Eleven v4 Turbo

Text to speech has been stuck in an awkward middle ground for a while. Output is usually intelligible, occasionally impressive in short bursts, but it rarely holds up across a full audiobook chapter, and it almost never sounds like someone who understood the sentence they were reading. ElevenLabs’ pitch with v4 is that the model treats a script the way a voice actor would: who is speaking, what just happened, how this particular line should land.

That framing is marketing, though the mechanics underneath are concrete. The model was rebuilt on a new architecture rather than tuned from v3, and the headline capability is direction tags written straight into the script. You can drop [laughs], [whispers], or [door slams] into the text and the model follows the sequence more reliably than v3 did, sound effects included. Whether that tag parsing holds up against messy real-world scripts is something only a real project will tell you.

Where v4 actually helps

The more interesting details solve long-session problems. Context stitching keeps pacing and delivery consistent regardless of script length, so a full book sounds like one take instead of a patchwork of separately generated chunks. Regeneration stability matters too: redoing a line once or fifty times should produce the same speaker, which is what you want when you’re fixing one bad sentence in a forty-minute narration and the voice can’t shift underneath you.

Professional Voice Clones return in v4 after sitting out v3, and they now work across the model’s full emotional range in every language they support. Voice cloning starts from ten seconds of audio, or you can train a professional clone for a closer match. Voice Design lets you describe a voice in a sentence and get it without a recording session, and the Pronunciation Dictionary handles the terms that usually break TTS: OAuth as oh-auth, Hülkenberg as hool-ken-berg, Reykjavik as rayk-yah-veek.

There’s a catalog angle as well. The voice library carries over 17,500 voices sorted into categories like Narration, Conversational, Social media, Character, Educational, Advertisement, Entertainment, and Multilingual, which helps if you’d rather cast than build.

Eleven v4 Turbo and the latency question

For anything conversational, the older bottleneck wasn’t expressiveness so much as waiting. Turbo claims median inference latency of roughly 100 milliseconds and median time to first speech of about 150 milliseconds, which puts it in the range where a voice agent stops feeling like a walkie-talkie exchange.

The streaming model is built for agent loops specifically. Push text as the LLM generates it and audio starts coming back before the sentence finishes. Turbo keeps v4’s expressive range, so confirmations, escalations, and holds land differently from one another instead of reading identically, and it handles multilingual speech with native-sounding accents in Japanese, Spanish, or Portuguese. Professional Voice Clones behave the same on both models, keeping a brand voice consistent from the first turn of a call to the last. Worth noting: those latency numbers are vendor medians, not guarantees, and your own infrastructure and region will move them.

Formats, pricing, and who this is for

Output matches every other ElevenLabs model. MP3 for podcasts and general listening, WAV or PCM for studio work and post-production, µ-law for telephony and call-center integrations. Sample rate and bitrate are set through the API, and a single generation handles up to 10,000 characters, with context stitching covering longer content.

Pricing stays freemium. The free tier gives 10,000 credits a month, roughly ten minutes of audio, which goes fast if you’re testing long-form material. Paid plans start at $6 a month with 30,000+ credits, professional voice cloning, and higher limits. Enterprise adds custom pricing, higher volume, SSO, and priority support. Compliance covers SOC 2 Type II, ISO 27001, PCI DSS Level 1, GDPR, and HIPAA-eligible workflows.

So who should care? Audiobook and podcast producers fighting voice drift across long sessions, game and animation teams needing character work, publishers converting journalism into audio, and developers wiring voice agents into customer-facing calls. The API offers REST and streaming endpoints with TypeScript and Python SDKs. If your use case is a thirty-second social clip, this is more model than you need, and cheaper tools will get you there.

Community ranking

With only one launch on the board yesterday, Eleven v4 and Eleven v4 Turbo takes the top spot by default. The reception in developer threads has been genuinely positive rather than merely uncontested, particularly around the return of Professional Voice Clones and the latency figures for Turbo. That said, the free tier’s ten-minute ceiling is a real constraint for anyone who wants to properly evaluate the model before committing, and the 10,000-character generation cap means long-form work still requires stitching logic on your end.

Quick links