Skip to content
English
FikirPilot content

Gemini 3.8 Live and 3.8 Live Extended Thinking unveiled

Updated: 15 Eyl 2026 · 3 min read · 428 words

Published: · Story reached us: · Processing time: 2 h 47 min

Gemini 3.8 Live and 3.8 Live Extended Thinking unveiled
A phone with a digital display sitting on a table

Google introduced the Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models, which aim to make voice interactions more natural, fluid and intelligent. The models offer complex reasoning, real-time visual context and the ability to perform tasks in the background without interrupting the conversation. Users can access the models today through the Gemini API, Google Workspace and the Gemini app.

Gemini 3.8 Live Extended Thinking ranked first overall in Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. In task completion, the model scored 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking test and 97.7% on Big Bench Audio. Gemini 3.8 Live placed second in Speech Agent Arena. In ServiceNow’s EVA-Bench evaluation, the models advanced the Pareto Frontier in complex workflows by balancing accuracy and conversation quality.

Gemini 3.8 Live processes visual inputs in near real time to add context to conversations and can automatically switch among 97 supported languages during a conversation. It can keep the conversation going while running tools and API calls in the background. Using visual context, it can guide employees through onboarding processes and play chess.

Gemini 3.8 Live Extended Thinking, meanwhile, thinks and speaks simultaneously, verbally communicating its progress through multi-step tasks. It can turn rough sketches and spoken feedback into functional React components, and can prepare reservations, asynchronous function calls, business plans and marketing tools through conversation.

The model is available in Docs Live, Gmail Live and Keep Live. Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents support the development of voice interfaces with the Gemini Live API. Google is also working with Salesforce, Genspark and Lumeris. All generated audio is marked with SynthID so it can be detected to counter misinformation.

Why it matters

This announcement offers a framework that moves voice AI beyond simple question-and-answer interactions by combining visual context, tool use and multi-step workflows within the same interaction. It therefore creates new application possibilities, particularly for Google Workspace users, developers and companies building conversational interfaces. The models’ results in various tests show that Google measures the balance between voice quality and task completion; however, it remains unclear to what extent these benchmarks will translate into everyday tasks such as making reservations, preparing business plans or providing compliance guidance. While marking audio with SynthID is intended to help distinguish generated content, how this will affect users’ trust decisions also remains an issue to watch.

Background

Gemini is not a new name in the FikirPilot archive: over the past 90 days, we have published 10 news stories mentioning the name, the latest dated September 8, 2026.

Source: Google DeepMind