Google Launches Gemini 3.8 Live Models to Cut Latency in Voice AI Interactions
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice processing models designed to handle real-time reasoning while maintaining natural conversation flow with users.

Google is rolling out two new voice-based AI models aimed at solving the responsiveness challenges that have plagued conversational AI systems. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking represent the company's most sophisticated voice processing offerings to date, featuring capabilities for near-instantaneous reasoning and the ability to process speech and thought simultaneously.
The models can trigger third-party software tool calls behind the scenes without interrupting the user experience, according to Google's announcement. This architecture allows AI agents to continue engaging naturally with users while executing assigned tasks in parallel, creating what Google describes as a more human-like interaction pattern without the jarring pauses users typically encounter when agents pause to search for information.
Benchmark Performance
Both models have achieved strong results across multiple evaluation frameworks. Gemini 3.8 Live Extended Thinking scored 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0. On the Speech Agent Arena benchmark, Gemini 3.8 Live placed second, while claiming first position on ServiceNow Inc.'s EVA-Bench. Additional performance metrics include 68.6% on T-Voice, 35.1% on T-Voice-banking and 97.7% on Big Bench Audio for the Extended Thinking variant.
Advanced Capabilities
https://x.com/GoogleAI/status/2099908000924193124?ref_src=twsrc%5Etfw
Both versions incorporate automatic language detection and mid-conversation language switching. The models understand and produce speech across 97 languages, support near real-time visual grounding, and employ early verbal acknowledgments such as "let me check that" to respond to user requests in a more conversational manner.

Google has also integrated SynthID, an invisible watermark embedded in generated audio files, to help identify synthetic content and combat misinformation.
Availability and Access
Gemini 3.8 Live is now accessible through the Gemini API, Google AI Studio, and as an enterprise private preview in Gemini Enterprise and Search Live. Gemini 3.8 Live Extended Thinking is available through those same channels plus Google Workspace integration in Docs, Gmail and Keep for subscribers, as well as the standalone Gemini Live applications.
Developers can embed both models into their own applications via Google's partner ecosystem, which includes Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents.
Pricing
https://www.youtube.com/embed/4P3iKlmQmM0?feature=oembed
The standard Gemini 3.8 Live model carries a cost of $0.005 per minute for audio inputs and $0.018 per minute for outputs. The Extended Thinking variant adds charges for reasoning tokens and supplementary inputs including video and documents.


