H
Howardism
Plate IIEntities中文HOWARDISM

Gemini 3.8 Live

Google's September 2026 live-dialogue pair — 3.8 Live (scale/cost) and 3.8 Live Extended Thinking (high-complexity), shipped as two models rather than a frontend/backend pair, with Extended Thinking claimed to 'reason and speak simultaneously'; the launch is argued almost entirely on third-party boards — Artificial Analysis Speech to Speech Index 82.6 vs GPT-Live-1 Astra's 81.5, τ-Voice 68.6%, Sierra τ³-Banking 35.1%, cost per hour of input audio $0.84 (3.8 Live) to $3.50 (Extended Thinking) — plus 97 languages with mid-conversation switching, near-real-time visual grounding, background tool execution and SynthID-watermarked audio; no price, no latency figure and no architecture disclosure anywhere in the post

Article metadata
Publication details
Published:September 21, 2026
Filed:Entity
Domain:Entities
Reading:7 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Gemini 3.8 Live

Sources#

What it is#

Two live-dialogue models announced together by Google's Gemini Audio Team (Tom Ouyang, Malini Jaganathan) on 2026-09-15 — Google's competitive answer to GPT-Live, arriving five days after OpenAI put GPT-Live-1 in the API.

Gemini 3.8 LiveGemini 3.8 Live Extended Thinking
Positioned for"scale and cost efficiency""high-complexity tasks"
Claimed propertyconversational intelligence, fluid dialogue, visual grounding"increased intelligence and multi-step reasoning"; "reasons and speaks simultaneously"
Consumer surfaceSearch LiveGemini Live, Workspace (Docs / Gmail / Keep)

Both ship the same day into the Gemini API and AI Studio for developers, and into private preview in Gemini Enterprise (with Gemini Enterprise for Customer Experience "coming soon"). Everything on this page is vendor-claim — a Google product announcement about Google's own models.

What Google claims it does#

  • Background tool execution while the conversation continues. "It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background." This is the fourth independent assertion of speak-while-thinking in this corpus, and like the previous three it arrives with no measurement — count and context on Interaction / Background Model Split.
  • Two named devices for covering the wait, both of which cut against the assertion above and are sold as features anyway: "early verbal cues like 'Let me check that…' to acknowledge prompts naturally," and "live progress narration to walk users through multi-step background tasks as they progress."
  • 97 languages with automatic mid-conversation switching — "automatically detects and transitions between 97 supported languages mid-conversation." No accuracy figure, no switching-latency figure, no named benchmark.
  • Near-real-time visual grounding — "processes visual inputs in near real-time, enriching conversations with context." Demonstrated only in untranscribed demo videos (live onboarding guidance, playing chess from the board).
  • SynthID watermarking on all generated audio — "woven directly into the audio output," the corpus's first provenance-marking claim on a live voice product.

The scoreboard#

The launch is argued almost entirely on third-party boards rather than internal evals — Artificial Analysis for three of the five charts, Sierra for one, ServiceNow for one — which is a different rhetorical footing from OpenAI's self-versus-self card set. All five charts are transcribed on Interactivity Benchmarks, which also carries what they can and cannot settle; the headline rows:

Chart (publisher)Gemini 3.8 Live ETGemini 3.8 LiveNearest competitor
Speech to Speech Index (Artificial Analysis)82.6% (High)76.0%GPT-Live-1 Astra 81.5% (Medium); Grok Voice Think Fast 2.0 81.3% (High)
Agentic Performance, τ-Voice (Artificial Analysis)68.6% (High)30.1%GPT-Live-1 Astra 67.9% (Medium)
τ³-Banking Leaderboard (Sierra)35.1% (High)—GPT-Live-1 Astra 32.0% (Medium)
Cost per hour of input audio, Big Bench Audio subset (Artificial Analysis)$3.50 (High)$0.84GPT-Live-1 Astra $5.83; Grok Voice Think Fast 2.0 $4.80

Plus two figures that appear in prose only, with no chart behind either: 97.7% on Big Bench Audio for Extended Thinking (the chart carrying that benchmark's name plots cost, not accuracy), and 3.8 Live "securing a second place in the Speech Agent Arena" — no score, no link, no date.

The interesting row is the one Google does not comment on. On the composite index the cheap tier beats the previous generation's thinking configuration (3.8 Live 76.0 vs Gemini 3.1 Flash Live High 71.5); on the agentic component the ordering inverts (30.1 vs 37.7). A single composite number and its own component disagree about which model is better, inside one vendor's chart set.

ServiceNow's EVA-Bench is the chart that does not flatter. Google's prose says "our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality." The dotted frontier drawn on the chart runs Gemini 3.8 Live → Extended Thinking (minimal) → GPT Realtime 2.1 → Scribe Realtime + GPT 5.4 + Eleven Flash v2, i.e. through a competitor and a third-party cascaded stack, and Extended Thinking (high) sits below it — dominated by GPT Realtime 2.1 on both axes. See Interactivity Benchmarks for the read and the transcription correction behind it.

What the post does not say#

Three absences, each load-bearing somewhere else in the wiki:

  • No price. Not per minute, not per token, not per audio hour. The only money on the page is a third party's measured cost, which is not a price list and does not separate a live layer from a backend. So the priced-interaction-boundary shape that GPT-Live-1 established on 2026-09-10 has no second vendor instance — see Interaction / Background Model Split.
  • No latency figure of any kind. Not turn-taking latency, not time-to-first-audio, not a delegation budget — remarkable for a launch whose partner quotes cite "impressive latency," and it means nothing here enters the cross-lab latency non-reconciliation on Interactivity Benchmarks.
  • No architecture. The post never says whether Extended Thinking is one model that thinks while it talks or a live model plus a delegated reasoner. The framing points one way and the vocabulary points the other; Interaction Models holds the argument.

Ecosystem#

Developer platforms named as building on the Gemini Live API: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents — "these platforms manage complex real-time media streaming infrastructure behind the scenes." (LiveKit also appears in NVIDIA's census of concurrent frontend/backend designs, via EXA Deep Researcher.) Partner companies named: Salesforce, Genspark, Lumeris, plus eight more appearing only as image-rendered testimonial cards that were not extractable at ingest.

Connections#

Sources#

  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, Tom Ouyang & Malini Jaganathan (Gemini Audio Team, Google), blog.google, published 2026-09-15 with an inline "Updated September 17, 2026" whose delta is unknown (vendor-claim, ~1,764 words). Five charts, all transcribed at ingest and all five images viewed during compile; the EVA-Bench scatter carries no data labels and its values are eyeballed against gridlines to ~1–2 points, so it is cited for ordering and dominance, never for a number. The DeepMind model card referenced by the post is not ingested.
§ end
Cited by 11
Related articles
  • Interactivity Benchmarks

    FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…

  • GPT-Live

    OpenAI's third-generation voice system (July 2026): a full-duplex voice model that listens and speaks simultaneously —…

  • Interaction / Background Model Split

    Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…

  • Live-Path Minimalism

    GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; deleg…

  • TML-Interaction-Small

    TML's first interaction model: 276B MoE / 12B active, audio+video+text in / text+audio out, 200ms micro-turns, async ba…