Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are two new text-to-speech models added to the Gemini family, described by Google as its most expressive audio generation models yet. The models generate custom character voices and direct scene dialogue across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Gemini 3.8 Flash TTS is built for deep creative direction and character design, while Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale. Google frames the release as a shift that transforms voice generation from static presets into a dynamic creative studio, serving creators, developers, and enterprises who want to produce richer, more expressive audio experiences.
The models complement the fast-growing Gemini Audio family, which already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. The problem the announcement frames is that voice generation has traditionally been locked to static presets, which limits how much a creator can shape a character, a narrator, or an agent's personality. Google states these models enable richer, more expressive audio experiences while also improving user experiences in products such as Gemini Notebook and Google Vids. The intended outcomes are high-quality audiobooks, podcasts, and real-time voice agents at scale, plus global dubbing and media localization with nuanced regional accents.
Creating and customizing voices is the first pillar of the release. Google says the models let teams scale up from 30 original voices to an infinite library, so an entirely original character voice or a consistent brand ambassador can be produced as needed. Generative voice design in Gemini 3.8 Flash TTS lets users create bespoke voices from scratch by customizing role, accent, and voice characteristics across more than 100 languages and dialects using natural language prompting; Google's examples include bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. Beyond generated voices, an expansive voice library offers 2,000+ production-ready voices with broad language coverage, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
Voice replication lets users recreate consistent vocal profiles from just a 30-second audio sample of their own voice or a voice they have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials intended to protect both developers and their vocal talent. A save-and-scale capability lets users save and manage the custom voices they designed, ensuring consistent performance and minimal drift across ongoing projects. Voice remixing is listed as coming soon: users will be able to pick a voice from the voice library and fine-tune timbre, pitch, pace, and accent, using prompts such as "add subtle Southern US accent" or "soften the delivery" to dial in specific characteristics.
Once a voice is chosen, both models provide precise control over how each line is delivered. Users can direct performance line by line by writing their own stage directions or letting Gemini steer delivery with natural script cues, covering everything from a calm customer service agent to a whispered suspense scene. Long-form generation maintains high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift, which Google calls ideal for podcasts and audiobooks. Native two-speaker scene staging directs multi-turn conversations seamlessly from a single script, whether for a podcast or dramatic storytelling, while keeping both voices distinctly separated with natural conversational turn-taking. Scripted vocal bursts and backchanneling add realistic conversational texture through non-verbal cues such as , , and , plus active-listening interjections like |mhm| or |yeah|, enabling precise comedic timing and reaction beats.
The overall approach is prompt-driven and script-driven: natural language prompts turn descriptions of role, accent, and character into vocal personas from scratch, and written scripts become fully performed dialogue scenes with directable delivery. Google positions these models as delivering expressive, high-quality speech generation built for global scale. Gemini 3.8 Flash TTS secured the number one overall spot on Hume AI's Voice Design Benchmark with a score of 71.4 and also led in accent modeling at 60.8. Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS took the number one and number two spots respectively on Hume AI's Overall Quality Index, and Google reports major improvements on use cases such as long-form content and dual-speaker screenplay control compared with Gemini 3.1 Flash TTS. In blind human preference evaluations on Voice Arena, the models secured top positions among competitors in key global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi, with support for over 100 languages.
For users, the stated benefits center on expressive, natural-sounding output that can be produced at scale without sacrificing reliability or consistency. Creators get granular creative control instead of preset-bound voices; developers get speech generation they can deploy into voice interfaces and voice agents; enterprises get multilingual reach and cost-efficient volume. Because the models maintain quality across hours of audio and keep two speakers clearly separated, teams can produce long-form and conversational content that holds together from beginning to end. Built-in safety tooling, including watermarking, is presented as a way to keep generated audio secure and to help prevent misinformation while still allowing ambitious creative work.
Concrete scenarios described in the announcement include high-quality audiobooks, podcasts, games, immersive audiobooks, interactive media, and real-time voice agents at scale. Developers can experience the speech generation capabilities in the Google AI Studio audio playground, which is built like a voice design workspace where they can prompt entirely new vocal identities from scratch or replicate their own voice and then bring them into a dual-speaker screenplay editor to direct line-by-line delivery. Google's partners are integrating the models to accelerate global dubbing, localize media with nuanced regional accents, and power conversational voice agents at scale. Example voices demonstrated in the announcement include a high-energy DJ voice from Melbourne, a super-tinny monotone robot voice, and a Japanese dragon brought to life.
Gemini 3.8 Flash TTS is rolling out for developers in the Gemini API and Google AI Studio, for enterprises coming soon via API in Gemini Enterprise, and for everyone in Gemini Notebook. Gemini 3.8 Flash-Lite TTS rolls out starting the same day for developers in the Gemini API and Google AI Studio, for enterprises coming soon via API in Gemini Enterprise, and for everyone in Google Vids. Through the Gemini API, developer platforms such as Agora, LiveKit, Pipecat, and Vercel enable developers to build and deploy high-performance speech generation experiences with ease. Google is also partnering with companies including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang. Voice replication through AI Studio is not available in Illinois, Texas, EEA, UK, Switzerland, and India. No pricing details are given in the announcement.
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS take text-to-speech beyond static presets by pairing natural-language voice design and voice replication with line-by-line performance direction, long-form stability, native two-speaker staging, and availability through Google AI Studio, the Gemini API, Gemini Notebook, and Google Vids. For anyone producing audio at any scale, the promise is expressive, natural-sounding speech with the control of a director and the safeguards of an enterprise platform.