Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models designed to give creators, developers and businesses greater control over AI-generated voices.

Users can create and direct custom voices

According to Google, the new models move beyond fixed synthetic voices by allowing users to create customised vocal identities using natural-language prompts. Users can specify characteristics such as accents, roles and other vocal attributes across more than 100 languages and dialects.

The higher-end Gemini 3.8 Flash TTS is aimed at applications including gaming, audiobooks, podcasts, dubbing and interactive media. It allows creators to direct how individual lines are delivered, including pacing, acting cues, dialect changes and conversational responses.

Google said users can access more than 2,000 production-ready voices, while custom voices can also be created and saved for repeated use.

Voice replication comes with safeguards

The model can recreate a vocal profile from a 30-second audio sample, provided the user has the necessary rights. Google said the process includes consent verification, while generated audio carries SynthID watermarking and C2PA credentials.

The Gemini 3.8 Flash-Lite TTS model is designed for high-volume applications where cost and scale are key considerations. Google highlighted potential uses such as large-scale dubbing, audio production and voice agents.

Both models support long-form audio and two-speaker conversations, with distinct voices and natural turn-taking. Non-verbal cues such as laughter, sighs and gasps can also be incorporated into scripted performances.

The models are available through Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. Google said platforms including Agora, LiveKit, Pipecat and Vercel are integrating the technology