Credit: Google “Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio,” Google wrote in its announcement. Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday. Both are rolling out from today in the Gemini API and Google AI Studio. They join the Gemini 3.8 family, which Google launched earlier this month.
Flash TTS is built for creative work such as games, audiobooks and podcasts. Flash-Lite TTS is the cheaper, high-volume option, aimed at dubbing and voice agents, according to Google. Voices designed from a prompt With Flash TTS, users can create a new voice from scratch by describing its role, accent and characteristics in plain language, across more than 100 languages and dialects. Google’s examples include a high-energy DJ from Melbourne and a Japanese dragon.
Users can also pick from a library of more than 2,000 ready-made voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English. Custom voices can be saved for reuse across projects. A voice remixing feature, which adjusts timbre, pitch, pace and accent, is coming soon. Both models let users direct delivery line by line with stage directions.
They support two-speaker scenes from a single script and long-form audio lasting hours. Scripts can include cues such as laughs, sighs and short interjections like “mhm”. Cloning a voice from 30 seconds Flash TTS can also recreate a voice from a 30-second audio sample. Google says users must supply a verbal consent recording from the voice owner, and the system checks that it matches the reference speaker before creating the voice.
Every clip generated by Google’s Gemini Audio models carries a SynthID watermark, according to the company. Replicated voices also carry C2PA content credentials. Benchmarks and partners Google says Flash TTS took first place on Hume AI’s Voice Design Benchmark, with a score of 71.4. It says Flash TTS and Flash-Lite TTS also ranked first and second on Hume AI’s Overall Quality Index.
In blind preference tests on Voice Arena, the models secured top positions in languages including Japanese, Hindi and Mexican Spanish, according to Google. The market already has specialists such as ElevenLabs and India’s Murf AI, and Adobe added speech generation to Firefly in August. Google named Figma, HeyGen, Wondercraft, Linguana, 99.co and Ollang as companies integrating the new models. Flash TTS is also coming to Gemini Notebook, and Flash-Lite TTS to Google Vids.
Access for enterprises through Gemini Enterprise is coming soon. Also tagged with













Leave a Reply