資料來源#
這是什麼#
Google 的 Gemini Audio Team(Tom Ouyang、Malini Jaganathan)於 2026-09-15 一同發布的兩款即時對話模型——這是 Google 對 GPT-Live 的競爭回應,時間比 OpenAI 將 GPT-Live-1 納入 API 晚五天。
| Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking | |
|---|---|---|
| 定位 | 「規模與成本效益」 | 「高複雜度任務」 |
| 宣稱特性 | 對話智慧、流暢對話、視覺 grounding | 「提升智慧與多步驟推理」;「一邊推理、一邊說話」 |
| 消費者產品介面 | Search Live | Gemini Live、Workspace(Docs / Gmail / Keep) |
兩款模型同日進入 Gemini API 與供開發者使用的 AI Studio,並在 Gemini Enterprise 開放私人預覽(Gemini Enterprise for Customer Experience 則「即將推出」)。本頁所有內容都是 vendor-claim——Google 對自家模型的產品公告。
Google 宣稱它能做什麼#
- 對話持續進行時,在背景執行工具。「它會在背景執行工具與 API 呼叫,同時繼續對話,因此模型可以回應請求並持續聊天,等待任務在背景完成。」這是本語料中第四次獨立宣稱能一邊說話一邊思考;和前三次一樣,沒有附上任何衡量數據——次數與脈絡見互動/背景模型分工。
- 兩種明確的等待期間安排,兩者都與上述宣稱相牴觸,卻仍被當成特色推銷:「像 『讓我查一下……』 這樣的初步口頭提示,自然地回應提示」,以及「即時進度旁白,隨著多步驟背景任務推進,引導使用者了解進展。」
- 支援 97 種語言,並能在對話中途自動切換——「自動偵測並在對話中途切換至 97 種支援語言。」沒有準確率、切換延遲數據或具名基準測試。
- 近乎即時的視覺 grounding——「以近乎即時的速度處理視覺輸入,為對話增添脈絡。」僅在未附逐字稿的示範影片中展示(即時引導新手操作、依棋盤下棋)。
- 所有生成音訊都加上 SynthID 浮水印——「直接融入音訊輸出」,是本語料中首個針對即時語音產品提出的來源標記宣稱。
排行榜#
這次發布幾乎完全以第三方排行榜而非內部評估來論述——五張圖表中有三張來自 Artificial Analysis,一張來自 Sierra,一張來自 ServiceNow——論述基礎不同於 OpenAI 自我對比的卡片組。五張圖表都轉錄在互動基準測試,其中也說明了這些圖表能判定與不能判定的事項;以下是主要數據列:
| 圖表(發布者) | Gemini 3.8 Live ET | Gemini 3.8 Live | 最接近的競爭者 |
|---|---|---|---|
| Speech to Speech Index(Artificial Analysis) | 82.6%(高) | 76.0% | GPT-Live-1 Astra 81.5%(中);Grok Voice Think Fast 2.0 81.3%(高) |
| Agentic Performance,τ-Voice(Artificial Analysis) | 68.6%(高) | 30.1% | GPT-Live-1 Astra 67.9%(中) |
| τ³-Banking Leaderboard(Sierra) | 35.1%(高) | — | GPT-Live-1 Astra 32.0%(中) |
| 輸入音訊每小時成本,Big Bench Audio 子集(Artificial Analysis) | $3.50(高) | $0.84 | GPT-Live-1 Astra $5.83;Grok Voice Think Fast 2.0 $4.80 |
另外有兩項數字只出現在正文中,背後都沒有圖表:Extended Thinking 在 Big Bench Audio 上得到 97.7%(以該基準名稱為標題的圖表呈現的是成本,不是準確率);以及 3.8 Live「在 Speech Agent Arena 拿下第二名」——沒有分數、連結或日期。
有意思的是 Google 沒有評論的那一列。在綜合指數上,平價版本勝過上一代的 thinking 配置(3.8 Live 76.0,Gemini 3.1 Flash Live High 71.5);但在 agentic 分項上,排名反轉(30.1 對 37.7)。同一組供應商圖表中,單一綜合分數與其中一個分項對哪款模型較好得出了相反結論。
**ServiceNow 的 EVA-Bench 是唯一不利於 Google 的圖表。**Google 的正文稱「我們的模型在複雜工作流程中推進 Pareto Frontier,成功兼顧準確度與對話品質。」圖表上的虛線前緣依序經過 Gemini 3.8 Live → Extended Thinking(minimal)→ GPT Realtime 2.1 → Scribe Realtime + GPT 5.4 + Eleven Flash v2,也就是經過一個競爭對手與一組第三方串接模型;而 Extended Thinking(high)落在前緣之下——在兩個座標軸上都被 GPT Realtime 2.1 壓過。圖表解讀與轉錄更正見互動基準測試。
文章未提及的內容#
以下三項缺漏,在 wiki 其他地方各有重要影響:
- **沒有價格。**沒有每分鐘、每 token 或每小時音訊價格。文中唯一的金額是第三方測得的成本,不是價目表,也沒有區分即時層與後端。因此,GPT-Live-1 在 2026-09-10 建立的「互動計價界線」形態,尚無第二家供應商的實例——參見互動/背景模型分工。
- **完全沒有延遲數據。**沒有輪替發言延遲、首次音訊時間或委派預算——對一場合作夥伴引述稱其「延遲表現令人印象深刻」的發布來說,這點相當值得注意;也代表沒有任何內容能納入互動基準測試的跨實驗室延遲資料調和工作。
- **沒有架構資訊。**文章從未說明 Extended Thinking 是能邊說邊思考的單一模型,還是即時模型加上委派推理器。宣傳說法指向一方,術語用法卻指向另一方;爭論見互動模型。
生態系#
文章提到以 Gemini Live API 為基礎建構服務的開發者平台:Agora、Fishjam、LangChain、LiveKit、Pipecat、Vercel、Vision Agents——「這些平台在幕後管理複雜的即時媒體串流基礎架構。」(LiveKit 也出現在 NVIDIA 對並行前端/後端設計的調查中,透過 EXA Deep Researcher。)文章提及的合作公司有:Salesforce、Genspark、Lumeris,另有八家只出現在以圖片呈現的推薦卡片中,匯入時無法擷取。
相關連結#
- Google DeepMind — 其研究實驗室;這是本語料中首個語音/即時互動相關資料
- GPT-Live — 直接競爭者,也是這些圖表所比較的模型
- 互動模型 — 其「互動性隨智慧提升」主張,可透過這次雙層級發布在同一供應商內進行檢驗
- 互動/背景模型分工 — 討論缺少的架構資訊,以及第四次「持續在場」宣稱之處
- 互動基準測試 — 五張圖表全數轉錄,並附跨供應商資料調和
- Full-Duplex Interaction — 將視覺 grounding 與對話中途切換語言列為宣稱的互動模式
- Artificial Analysis — 五張圖表中三張的評估方
- Cost-per-Task Over Cost-per-Token — 將輸入音訊每小時成本作為計費分母
- Large-Scale Test-Time Compute — 「Extended Thinking」是在即時路徑上設定的測試時運算調節項;圖表為每根長條都標示了運算強度設定
資料來源#
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, Tom Ouyang & Malini Jaganathan (Gemini Audio Team, Google),
blog.google,於 2026-09-15 發布,並附有「Updated September 17, 2026」的內文註記,但更新內容不明(vendor-claim,約 1,764 字)。五張圖表皆於匯入時轉錄,編輯時也逐一檢視;EVA-Bench 散佈圖沒有資料標籤,數值是依格線目測估算,誤差約 1–2 點,因此只用來判斷排序與支配關係,絕不引用為數值。文章提及的 DeepMind 模型卡未匯入。
Cited by 11
- Full-Duplex Interaction×4
One boundary note for the section below on speech-to-speech. The survey's own census (Table 1, 43…
- Artificial Analysis×3
Agentic Performance (τ-Voice) · Interactivity Benchmarks, Gemini 3 8 Live · AA's run of the τ-Voice…
- Interaction Models×3
Gemini 3 8 Live — the fourth live-voice product, and the first whose public account describes no…
- Google DeepMind×2
Gemini 3 8 Live — the lab's live-dialogue pair, and its entry point into the interaction/multimodal…
- GPT-Live×2
Gemini 3 8 Live — the competitor that arrived five days later and scored it on four third-party…
- Interaction / Background Model Split×2
Gemini 3 8 Live — the fourth live-voice product: no backend named, no price on either half, and…
- Interactivity Benchmarks×2
Google's Gemini 3.8 Live launch (2026-09-15, vendor-claim) arrives five days after OpenAI's and…
- Content-Driven Intervention
Gemini 3 8 Live — the second shipping voice product to claim nothing on this axis, and the one that…
- Cost-per-Task Over Cost-per-Token
Gemini 3 8 Live — the second frontier voice launch of the month, which publishes no price at all…
- Entities — People, Orgs, Tools & Projects
Gemini 3 8 Live — Entity. Google's September 2026 live-dialogue pair — 3.8 Live (scale/cost) and…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
Nothing from the interaction-model line appears at all: no Tml Interaction Small, no Inkling, no…
Related articles
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
- GPT-Live
OpenAI's third-generation voice system (July 2026): a full-duplex voice model that listens and speaks simultaneously —…
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Live-Path Minimalism
GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; deleg…
- TML-Interaction-Small
TML's first interaction model: 276B MoE / 12B active, audio+video+text in / text+audio out, 200ms micro-turns, async ba…
