- Google is rolling out Gemini 3.8 Flash TTS and Flash-Lite TTS, its latest text-to-speech models for more expressive AI-generated audio.
- Gemini 3.8 Flash TTS can create custom voices, replicate an authorized voice from a 30-second sample, and control accents, pacing, emotion, and two-speaker dialogue.
- The new models are reaching Gemini Notebook and Google Vids, with both also available to developers through Google AI Studio and the Gemini API.
AI voiceovers can read a script. But getting them to actually perform it with the right accent, pacing, emotion, pauses, and even the occasional sigh is another story. To address this, Google has spent weeks growing Gemini 3.8 from the original Gemini 3.8 Flash into new Live and Live Extended Thinking models. And now it wants the same family to act the part, with Gemini 3.8 Flash Text-To-Speech (TTS) and Flash-Lite TTS rolling out.
The biggest change for users is control. In a blog post, Google says Gemini 3.8 Flash TTS can create an original voice from a natural-language description, with support for more than 100 languages and dialects. You also get over 2,000 production-ready voices, and voice replication can recreate a consistent voice from a 30-second sample, provided you have permission to use it.
​Â