Models

Google Rolls Out Dual Text-to-Speech Models With Language and Quality Tiers

Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS through its cloud platform, offering developers a choice between superior audio fidelity and cost-efficient performance across different language support levels.

·3 min read
Google launches two benchmark-topping speech generation models
Google launches two benchmark-topping speech generation models

Google LLC unveiled two fresh text-to-speech generation models on its cloud infrastructure: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The pair share comparable application programming interfaces, allowing developers to deploy them together without significant friction. Flash-Lite TTS prioritizes affordability and rapid inference, whereas Flash TTS delivers enhanced audio quality at a premium cost point. The company sees potential applications ranging from audiobook production to other audio content creation.

A key distinction separates the two offerings in terms of linguistic reach. Flash TTS launches with the ability to synthesize speech across 130 languages, while Flash-Lite TTS covers 101 languages at launch.

Voice Customization and Personalization

Both models grant developers access to a collection exceeding 2,000 pre-built voices. Beyond selecting from this library, developers can construct personalized voices using natural language descriptions. Google permits fine-tuning of characteristics including vocal timbre, regional accent and speech rhythm.

An alternative customization path involves generating a synthetic voice from a 30-second audio recording. Google mandates that developers obtain explicit consent from the original speaker before creating a voice based on their sample. The company has signaled plans to introduce a third customization method in the future, enabling voice creation through modification of existing prepackaged voices.

The models also enable developers to shape how an AI speaker delivers content. Google reports that developers can embed oratory instructions into each script line, allowing the models to produce audio features such as non-lexical vocalizations and changes in pacing.

Watermarking and Content Attribution

Google employs a technique named SynthID to insert an imperceptible audio watermark into synthesized speech. While undetectable to human ears, AI-based detection systems can identify this marker. Additionally, Google appends a C2PA record to each generated audio file, documenting generation timestamps, modification history and other relevant metadata.

Benchmark Performance

Google assessed both models using an audio quality benchmark created by Hume AI Inc. Flash TTS secured first place and Flash-Lite TTS placed second in the evaluation. The models also demonstrated superior performance relative to competing algorithms on multiple language-specific iterations of Voice Arena, a benchmark that rates text-to-speech output quality using human evaluations.

These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids

Google staffers Leland Rechis and Alan Cowen

The two new models represent part of Google's expanding suite of audio processing algorithms. The company has previously introduced models designed for voice agents, transcription and translation applications.