# Gemini 3.8 Live: Google's voice AI that thinks while talking

> Google launched voice AI models that reason and run tasks mid-conversation.

*Google's new speech-to-speech models can pause, reason, run tools and pick back up without leaving you hanging on the line.*

By Behzad Hosseini · SuggestedTech
Canonical: https://suggestedtech.com/news/gemini-3-8-live-google-s-voice-ai-that-thinks-while-talking

Google has released two new voice AI models built to sound less robotic when they have to stop and think. Instead of going quiet while working something out, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking keep talking, letting you know they're on it, then coming back with an answer once they've done the work.

## What happened

Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026. Both are "speech-to-speech" models, meaning they take spoken audio in and give spoken audio back, rather than converting your voice to text first.

The standard version, 3.8 Live, is built for speed and lower cost at large scale. The Extended Thinking version can do multi-step reasoning while it speaks, with a choice of low, medium or high reasoning effort. Rather than falling silent to work through a problem, it acknowledges you out loud, keeps the conversation flowing, runs any tools it needs in the background, and reports back once it has an answer.

On the Artificial Analysis Speech to Speech Index, Extended Thinking (High) launched in first place with a score of 82.6, ahead of OpenAI's GPT-Live-1 (Astra, medium) at 81.5 and xAI's Grok Voice Think Fast 2.0 High at 81.3. It also scored 68.6% on the Tau-Voice benchmark, compared with 30.1% for the standard 3.8 Live, and 97.7% on Big Bench Audio.

Speed has improved too, according to [MarkTechPost](https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/): time to first audio response is 1.18 seconds for Live and 1.35 seconds for Extended Thinking (High), down from 2.99 seconds on the previous Gemini 3.1 Flash Live model, as reported by [TechRepublic](https://www.techrepublic.com/article/news-gemini-3-8-live-models/).

## What it means for you

If you use voice assistants or customer-service bots, this points to less awkward waiting and more capable responses. The models can switch between 97 languages mid-call, so a single conversation could shift language without restarting.

Every response carries a SynthID watermark, a way of marking audio as AI-generated.

Extended Thinking is coming to Gemini Live, Docs, Gmail and Keep for subscribers, so people using Google's own apps should see it appear there rather than needing a separate tool.

## How to try it

Both models are generally available now, hosted only, through the Gemini API and Google AI Studio.

- Audio input costs $0.005 a minute; audio output costs $0.018 a minute
- Developer platforms including Agora, LangChain, LiveKit, Pipecat and Vercel already support the new models
- Consumers using Gemini Live, Docs, Gmail or Keep will get Extended Thinking as it rolls out to those apps

## Key takeaways

- Gemini 3.8 Live talks back in about 1.2 seconds instead of going silent to think
- The Extended Thinking version topped a major speech-to-speech leaderboard at launch
- It's available now via the Gemini API and Google AI Studio, priced by the minute

## Sources

- [Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents](https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/) — MarkTechPost, 2026-09-15
- [Google Launches Gemini 3.8 Live Models That Can Reason While They Talk](https://www.techrepublic.com/article/news-gemini-3-8-live-models/) — TechRepublic, 2026-09
