資料來源#
- Build more natural voice experiences with GPT‑Live‑1 in the API
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
這是什麼#
一家獨立的 AI 基準測試機構,發布模型排行榜與衍生指數。它是本維基中引用最多的第三方評估機構,也是來源溯源鏈最薄弱的一家:截至 2026-09-21,沒有任何 Artificial Analysis 出版物被匯入 raw/,此後也沒有——每個透過產品公告進入本維基的 AA 數字,都是經由廠商轉述的二手資料——可能出自系統卡、技術報告,或重現 AA 圖表的發布文章。這點值得明白說出來,因為由受評廠商引用的第三方排行榜,與直接閱讀原始資料中的第三方排行榜,並非同一種證據。
曾直接閱讀原始資料(2026-09-22)#
截至 2026-09-21,沒有任何 Artificial Analysis 出版物被匯入 (已於 2026-09-22 更新。) Zhu(Oxford Internet Institute,arXiv 2608.29420,raw/。empirical,單一作者預印本)是本語料中第一個以 AA 排行榜為主要資料集的來源,而且來自學術第三方,而非重現圖表的廠商。它仍不是 AA 的出版物——AA 在此尚未發布任何資料——但數字是從 AA 自己的頁面讀取,而不是從系統卡讀取;這與下表所有資料的來源類別不同。
這類來源才能提供的該組織四項資訊:
- 公開頁面是唯一可存取的介面。「Data API 拒絕了未經驗證的請求,因此我們擷取公開頁面並儲存確切位元組。」此快照以 SHA-256 固定雜湊值(
6f19f8f0…,收錄於隨附的 manifest),擷取日期為 2026-07-06。任何想取得可重現 AA 數字的人,都得擷取 HTML,而非呼叫 API。 - **排行榜規模遠超本維基過去接觸的排行榜。**共 548 種模型組態,附有各基準測試的分數、發布日期與供應商中繼資料,涵蓋至少十四項基準測試;同一個基礎模型可能以數種推理與運算量設定出現,因此「組態」與「模型」是不同的計數(421 種組態縮減為 89 個具備完整十二項基準測試組合的不同基礎模型)。
- **AA 在同一套 harness 下自行執行所有基準測試,並建構其中三項。**研究中的十二項基準測試全由營運方在單一 harness 下執行,而 AA-Omniscience、AA-LCR 與 Terminal-Bench Hard 由 AA 自行建構或選取子集。這項特性使其資料網格獨具價值(完全沒有廠商自我報告,正是公開分數矩陣常見的混淆因素——見 Benchmark Score Redundancy),同時也限制了它(營運方特有的題目建構,會與人們在其上測量的結構混淆)。
- **實際上,其 Intelligence Index 是能力層級標記。**在因子空間中對 409 個模型進行分群,得到兩個依發布日期區分的群集:302 個較舊模型的 Intelligence Index 平均值為 13.7(發布日期中位數為 2025-09),107 個較新模型則為 37.1(發布日期中位數為 2026-03)。
稽核對排行榜結構的結論,一句話說完:AA 的十二項基準測試表現出近乎單一維度的組合(非對角線 Spearman ρ 平均值 = 0.79,保留一個因子,解釋 74.5% 的共同變異),其主軸與發布日期的關聯為 R² = 0.505——因此,AA 對相隔數月發布模型所做的排名,很大程度上就是依發布日期排序。完整分析見 Economic Benchmark Construct Validity。
本語料接觸過的排行榜#
| 排行榜 | 出現處 | 備註 |
|---|---|---|
| Intelligence Index | Deep Research Agents | 用作能力排序;代理程式表現明確呈現非單調性 |
| GDPval-AA(Elo,v1 與 v2) | GDPval Benchmark, Claude Opus 4.8, Kimi (Moonshot AI), The Open-Weight Frontier Gap | 以 OpenAI 的 GDPval 為基礎建立的 Elo 排行榜;v1 與 v2 無法比較 |
| AA-Briefcase(Elo) | Kimi (Moonshot AI), The Open-Weight Frontier Gap | 長期任務代理程式 Elo 排行榜 |
| Conversational Dynamics | Interactivity Benchmarks, GPT-Live | 接近飽和:三代 OpenAI 模型依序為 95.3 → 95.7 → 97.3 |
| Speech to Speech Index | Interactivity Benchmarks, Gemini 3.8 Live | 本語料首個跨廠商語音指數(2026-09) |
| Agentic Performance (τ-Voice) | Interactivity Benchmarks, Gemini 3.8 Live | AA 執行 τ-Voice 任務系列的結果 |
| 每小時輸入音訊成本 | Interactivity Benchmarks, Cost-per-Task Over Cost-per-Token | 在工作負載上實測的成本,不是價目表——一種少見而實用的呈現方式 |
目前記錄的兩項風險#
**版本變更會在不知不覺中破壞可比性。**AA 重新評分了 GDPval-AA,而本維基曾將 Opus 4.8 系統卡中的 v1 數字,與後續系統卡中的 v2 數字並列引用,彷彿它們屬於同一個排行榜。完整分析見 GDPval Benchmark。一般而言,第三方維護的衍生排行榜可能在名稱不變的情況下修訂;引用它的廠商標示的是自家發布日期,而不是排行榜的日期。
第三方署名,第一方方法。Google 的 Gemini 3.8 Live 發布文章中的五張圖表,標示的發布者分別是 Artificial Analysis(三張)、Sierra(一張)與 ServiceNow(一張),但五張圖表都附上相同的方法註腳——連到 deepmind.google/models/evals-methodology/gemini-3-8-live,也就是廠商自己的頁面,而非 AA 的頁面。分數很可能來自 AA;但讀者被引導查閱的測試流程則屬於 Google。沒有任何資訊能確認由哪一方執行哪些模型,以及採用何種運算量設定;圖表中的運算量標籤(自家模型為 High,最接近的競爭者為 Medium)也沒有任何說明。
為何這對本維基很重要#
AA 在本語料中扮演其他評估機構都沒有的角色:它是競爭廠商都會引用的中立計分板,也是兩家實驗室的數字在此對上的唯一途徑。最明確的例子是 Sierra 的 τ³-Banking 排行榜:OpenAI 自行發布的系統卡與 Google 的競品圖表,都列出 gpt-live-1 + Astra 32.0%、gpt-realtime-2 10.3%——兩個立場相互對立的參與方列出完全相同的數據,是本語料首度出現逐位數字一致的跨廠商佐證(Interactivity Benchmarks)。同一節也記錄了相反案例:兩家廠商以不同名稱呈報相同組態,分別得到 67.9% 與 86.2%;因此,只有在排行榜確實共用時,共用排行榜才能提供佐證。
相關連結#
- Economic Benchmark Construct Validity — 本語料中唯一對 AA 排行榜進行的外部稽核,也是首次直接讀取 AA 資料:固定雜湊值的 2026-07-06 快照,涵蓋 548 種組態;將十二個排行榜視為心理計量題項,發現一個與發布日期相關的主導因子
- Benchmark Score Redundancy — 為何 AA 這種單一 harness、沒有廠商自我報告的資料網格,正是該領域人人想要、卻無人擁有的分數矩陣
- Interactivity Benchmarks — 彙整、加權並調和語音領域排行榜之處
- GDPval Benchmark — Elo 排行榜底層的主要基準測試,以及 v1/v2 不可比較一事
- Measuring Beyond Accuracy Saturation — 準確率不再能區分模型時,本維基轉而採用 Elo 排行榜;其中多數由 AA 提供
- Compute-Controlled Benchmarking — AA 圖表會標示每個長條的推理運算量設定,揭露的運算資訊多於多數資料網格,但仍不足以構成受控比較
- Gemini 3.8 Live, GPT-Live — 發布時以這些排行榜作為論據的兩項語音產品
- Kimi (Moonshot AI), Claude Opus 4.8, The Open-Weight Frontier Gap — 引用廠商系統卡中 Elo 排行榜的模型頁面
資料來源#
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google,2026-09-15(
vendor-claim):三張以 AA 為標題的圖表(Speech to Speech Index、Agentic Performance τ-Voice、Cost per Hour of Input Audio),由廠商重現,並附上廠商的方法註腳 - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI,2026-09-10(
vendor-claim):AA Conversational Dynamics 卡片 - One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu(Oxford Internet Institute,單一作者,未經同儕審查的預印本),arXiv 2608.29420,2026-08-29,
empirical:固定雜湊值的 2026-07-06 排行榜快照、遭拒的 Data API 請求、548 種組態的規模、單一 harness 與營運方建構基準測試的事實(附錄 F),以及 Intelligence Index 的層級分群(附錄 K)。這不是 AA 出版物,而是學術第三方閱讀 AA 公開頁面的研究。 raw/中沒有任何 Artificial Analysis 出版物。上方的排行榜說明是根據重現排行榜的廠商文件,以及已整理這些排行榜的維基頁面彙整而成;對本維基而言,此處列出的每個排行榜,其組成、加權方式與執行流程都未知。
Cited by 10
- Economic Benchmark Construct Validity×3
Artificial Analysis — the operator whose board is the dataset here, and the first source in this…
- Cost-per-Task Over Cost-per-Token×2
And five days later a third party measures the unit on a workload (Google, 2026-09-15). Google's…
- GDPval Benchmark×2
The caveat that bites this page specifically. The GDPval column in that snapshot is the Artificial…
- Gemini 3.8 Live×2
The launch is argued almost entirely on third-party boards rather than internal evals — Artificial…
- Google DeepMind×2
The disclosure posture inverts again, and this time in the lab's favour on one axis and against it…
- Interactivity Benchmarks×2
Artificial Analysis — the evaluator behind the Conversational Dynamics card, the Speech to Speech…
- Benchmark Score Redundancy
Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author preprint) runs the…
- Compute-Controlled Benchmarking
Everything on this page treats compute as the missing control. Zhu (Oxford Internet Institute,…
- GPT-Live
Artificial Analysis — the evaluator behind its Conversational Dynamics card and behind three of the…
- Entities — People, Orgs, Tools & Projects
Artificial Analysis — Entity. The third-party evaluator this corpus quotes most and has almost…
Related articles
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Cost-per-Task Over Cost-per-Token
Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
