Google recently released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, introducing near real-time reasoning, synchronized voice and thinking processing, as well as background tool and API call capabilities. The two models have been gradually launched on developer platforms such as Gemini API and Google AI Studio, and are respectively applied to scenarios including Search Live, Gemini Enterprise and Google Workspace, Gemini Live, etc.

A core capability of the two models is supporting AI to perform background tool and API calls during continuous voice conversations, so users do not need to pause communication for network searches or task execution. The models also support automatic detection of 97 languages and language switching during conversations, near real-time visual positioning, and can inform users of the current status through language prompts such as "Let me check...".
In terms of performance, Gemini 3.8 Live Extended Thinking scored 82.6 in the Artificial Analysis Speech to Speech Quality Index; Gemini 3.8 Live ranked second in the Speech Agent Arena. The Extended Thinking version scored 68.6% in the T-Voice test, 35.1% in the Sierra T-Voice benchmark test, and reached 97.7% in the Big Bench Audio test.
In addition, Google has partnered with platforms such as Agora, LiveKit, Vercel, Pipecat, Fishjam, and Vision Agents, making it easier for developers to integrate the new models. Google also stated that all audio generated by the two models includes an invisible SynthID watermark for identifying AI-generated audio.
This upgrade further expands real-time voice interaction from "listening and speaking" to understanding, reasoning, and task execution, promoting the development of voice AI towards continuous, multi-task intelligent agent interactions.
Join Now