Google’s new speech model Gemini 3.8 Live supports real-time reasoning
Google LLC is trying to address the latency problem associated with voice-based artificial intelligence agents with the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking today.
They’re billed as the company’s most advanced voice processing models released so far, and they’re capable of near-real-time reasoning and simultaneous speech-and-thought processing. They can also execute third-party software tool calls in the background, Google said in a blog post.
The new models boast top-tier benchmark results, if Google is to be believed, with Gemini 3.8 Live Extended Thinking achieving a brand new high score of 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0. Gemini 3.8 Live was ranked second place on the alternative Speech Agent Arena benchmark and first on ServiceNow Inc.’s EVA-Bench. And Gemini 3.8 Live Extended Thinking achieved a score of 68.6% on T-Voice, 35.1% on T-Voice-banking and 97.7% on Big Bench Audio.
Introducing our most advanced Gemini Audio models yet 🗣
Gemini 3.8 Live and 3.8 Live Extended Thinking let you speak, collaborate, and execute tasks seamlessly, meaning conversing with AI just got a lot more natural.
So, what’s the difference between these two models? Let’s… pic.twitter.com/Dfy7j4zxOh
— Google AI (@GoogleAI) September 15, 2026
Google explained that the new models are designed to execute tool and application programming interface calls in the background, while maintaining a natural pace of conversation with human users. This means that AI agents can continue chatting with their agents while simultaneously working on tasks that were just assigned to it. According to Google, it provides a more natural, human-like calling experience, without users seeing constant interruptions while their agents run off to browse the internet in search of answers to something.
![]()
There are some advanced features in these models too, Google said. For instance, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking both support automatic language detection and have the ability to switch language in mid-conversation. They’re able to understand and generate speech in 97 languages, they support near real-time visual grounding, and they can also use early verbal cues such as “let me check that,” to acknowledge user’s prompts in a more natural, lifelike way.
Google said Gemini 3.8 Live is available now via the Gemini API and Google AI Studio, and can also be accessed as an enterprise private preview in Gemini Enterprise and Search Live. As for Gemini 3.8 Live Extended Thinking, it’s available through the same channels and also in Google Workspace via Docs, Gmail and Keep (for subscribers) and in the Gemini Live applications.
Developers will also be able to integrate the models within their applications through Google partner platforms such as Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents, Google said. When audio files are generated, they’ll carry an invisible SynthID watermark, which can be used to detect misinformation, Google added.
With regard to pricing, Google said the standard Gemini 3.8 Live model costs $0.005 per minute for audio inputs and $0.018 per minute for outputs. The Extended Thinking model also charges for reasoning tokens and for additional inputs such as video and documents.
Images: Google
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.