Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, its most advanced live dialogue models to date . These models enhance voice interactions by integrating complex reasoning, real-time visual context, and background task execution for more natural, fluid conversations.
The introduction of these models aims to make conversing with AI feel more intuitive and intelligent, especially for complex tasks . They represent a significant step in developing voice agents that can reason and execute tools without breaking conversational flow.
Advancing Real-Time Voice Interaction
Google's Gemini 3.8 Live and Extended Thinking models redefine voice AI by enabling fluid, intelligent conversations that handle complex reasoning and visual context in real-time. They execute background tasks without interrupting dialogue, making voice agents more collaborative and intuitive .The models are designed to power AI-driven reasoning and agentic workflows, eliminating awkward silences during interactions . Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with immediate context. It also detects and transitions between 97 supported languages mid-conversation.
For deeper reasoning tasks, 3.8 Live Extended Thinking simultaneously reasons and speaks, providing early verbal cues like, “Let me check that...” to maintain conversational flow. This allows the model to acknowledge requests and continue chatting while tasks complete in the background .
How Do These Models Perform?
Gemini 3.8 Live Extended Thinking achieved the #1 spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, demonstrating superior performance in agentic task completion and reasoning. The base Gemini 3.8 Live model also ranked highly in user preference .Specifically, Gemini 3.8 Live Extended Thinking scored 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark for agentic task completion. It also showed strong reasoning capabilities with 97.7% on Big Bench Audio .
Model Feature | Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking |
|---|---|---|
Primary Focus | Scale, cost efficiency, fluid dialogue | High-complexity tasks, multi-step reasoning |
Visual Input | Near real-time processing | Near real-time processing |
Language Support | 97 languages (mid-conversation transition) | 97 languages (mid-conversation transition) |
Top Benchmark Score | Second in Speech Agent Arena | #1 on Artificial Analysis' Speech to Speech Quality Index (82.6) |
Both models are available to developers through the Gemini API and Google AI Studio . They are also rolling out across Google's consumer products, including Gemini Live, Google Workspace (Docs Live, Gmail Live, Keep Live), and Search Live.








