Google introduces Gemini 3.8 Flash and Flash-Lite text-to-speech models

Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, text-to-speech models that expand its voice library from 30 to over 2,000 voices across more than 100 languages.

Google introduces Gemini 3.8 Flash and Flash-Lite text-to-speech models

Google has launched two new text-to-speech models within its Gemini family, shifting voice generation from static presets to a dynamic creative tool. The models-Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS-give creators, developers, and enterprises granular control over voice design, performance direction, and high-volume audio production across more than 100 languages.

Two models built for different workflows

Gemini 3.8 Flash TTS focuses on deep creative direction and character design. Users can build entirely new voices from scratch using natural language prompts, specifying role, accent, and vocal characteristics. The model supports line-by-line performance control with acting cues, pacing adjustments, dialect shifts, and backchanneling-non-verbal sounds like laughter or sighs that add conversational texture.

Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient production. It is optimized for dubbing, audio content creation, and expressive voice agents, offering fine-grained control over tone, pacing, and expressive nuance without the full creative toolset of its larger counterpart.

Voice creation and customization

The 3.8 Flash TTS model expands Google's voice library from 30 original voices to over 2,000 production-ready options, including regional varieties such as Mexican Spanish, Quebec French, and Scots English. Users can also replicate a voice from a 30-second audio sample, provided they have the rights to use it. Google has built in consent verification, requiring a verbal confirmation from the voice owner that matches the reference speaker before replication proceeds.

Generated audio carries SynthID watermarking and C2PA credentials. A voice remixing feature, which will let users adjust timbre, pitch, pace, and accent on existing library voices, is listed as coming soon.

Performance control and scene staging

Both models allow users to direct vocal delivery line by line through script cues or explicit stage directions. The system supports native two-speaker scene staging from a single script, keeping voices distinct during multi-turn conversations. This suits podcast production, dramatic storytelling, and interactive voice agents. Long-form generation maintains voice quality and pacing across hours of audio with minimal speaker drift.

On benchmarks, Gemini 3.8 Flash TTS placed first on Hume AI's Voice Design Benchmark with a score of 71.4 and led in accent modeling at 60.8. The Flash and Flash-Lite models took the top two spots on Hume AI's Overall Quality Index. In blind human evaluations on Voice Arena, both models ranked among the top competitors in languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Availability and integrations

Both models are rolling out today in the Gemini API and Google AI Studio. Enterprise access via Gemini Enterprise is coming soon. Consumers will find Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are enabling deployment through the Gemini API. Companies such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are integrating the models for dubbing, media localization, and voice agents.

Why this matters for creatives, writers, and customer support teams

These models shift voice production from a fixed menu of presets to a prompt-driven creative workflow. For writers and audio producers, the ability to direct character dialogue with stage directions and backchanneling means scripts can carry performance intent directly into production. Customer support teams building voice agents gain finer control over tone and pacing without sacrificing cost efficiency at scale. The consent verification and watermarking system also provides a practical framework for voice talent protection, which matters for brands and agencies managing vocal IP across campaigns. Professionals looking to build these skills can explore Text to Speech AI Courses and AI Voice Modulation Courses to apply these capabilities in their workflows.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)