On September 23, 2026, Google released two new speech synthesis (text-to-speech, or TTS) models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Google positions the two as its most expressive speech generation models.
How the two models differ
- Gemini 3.8 Flash TTS: For uses where expressiveness matters, such as fine-grained direction and character creation
- Gemini 3.8 Flash-Lite TTS: For generating large volumes of speech at low cost
Key features
- More than 2,000 ready-to-use voices
- Support for more than 100 languages and dialects, with new voices designed simply by describing the role, accent and vocal characteristics in text
- Recreation of a voice from a 30-second audio sample
- Line-by-line control over speaking pace, emotion and even the natural sounds of real conversation
- A consistent voice across long passages, plus generation of conversations between two speakers
According to Google, Flash TTS ranked first (with a score of 71.4) on Hume AI's voice design benchmark. On the Voice Arena leaderboard, both models also achieved top-tier rankings ahead of competitors in multiple languages, including Japanese, Brazilian Portuguese and Vietnamese.
Safeguards
All generated audio carries SynthID, a digital watermark indicating that it was made by AI. Voice cloning includes a mechanism to verify the speaker's consent, and the models also support C2PA for recording provenance.
Voice cloning is not available in the US states of Illinois and Texas, the European Economic Area (EEA), the UK, Switzerland or India.
Where to use it
Both models are available in the Gemini API and Google AI Studio, with support in Gemini Enterprise coming soon. Flash TTS is also available to all users in Gemini Notebook, and Flash-Lite TTS in Google Vids.
CoAI's take
Beyond the sheer number of voices, a standout feature is the ability to create new voices just by giving text instructions. The models look set to become an option for a wide range of audio work, from narration and video production to building voice assistants. If you plan to use voice cloning, check in advance where it is available and how the consent verification process works.