H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Artificial Analysis

本語料引用最多、卻幾乎從未直接閱讀的第三方評估機構——其 Intelligence Index、GDPval-AA Elo board、AA-Briefcase、Conversational Dynamics、Speech to Speech Index、τ-Voice 與 cost-per-hour-of-input-audio 圖表,都是透過廠商系統卡與發布文章才以二手資料進入本維基;目前記錄了兩項風險:一次未公告的重新評分,導致 GDPval-AA v1 與 v2 無法比較;以及一組 Google 圖表將排行榜歸於 Artificial Analysis,註腳卻指向廠商自己的方法頁面。截至 2026-09-22,首次直接閱讀原始資料的第三方研究——Oxford 對一份固定雜湊值的快照進行稽核,發現單一因子解釋了超過 74.5% 的共同變異,且與發布日期的關聯為 R² = 0.505——使這項來源溯源缺口縮小了一半

Article metadata
Publication details
Published:September 21, 2026
Filed:Entity
Domain:Entities
Tags:Type/entityLLM Evaluation
Reading:9 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Artificial Analysis 插圖

資料來源#

這是什麼#

一家獨立的 AI 基準測試機構,發布模型排行榜與衍生指數。它是本維基中引用最多的第三方評估機構,也是來源溯源鏈最薄弱的一家:截至 2026-09-21,沒有任何 Artificial Analysis 出版物被匯入 raw/,此後也沒有——每個透過產品公告進入本維基的 AA 數字,都是經由廠商轉述的二手資料——可能出自系統卡、技術報告,或重現 AA 圖表的發布文章。這點值得明白說出來,因為由受評廠商引用的第三方排行榜,與直接閱讀原始資料中的第三方排行榜,並非同一種證據。

曾直接閱讀原始資料(2026-09-22)#

截至 2026-09-21,沒有任何 Artificial Analysis 出版物被匯入 raw/。 (已於 2026-09-22 更新。) Zhu(Oxford Internet Institute,arXiv 2608.29420,empirical,單一作者預印本)是本語料中第一個以 AA 排行榜為主要資料集的來源,而且來自學術第三方,而非重現圖表的廠商。它仍不是 AA 的出版物——AA 在此尚未發布任何資料——但數字是從 AA 自己的頁面讀取,而不是從系統卡讀取;這與下表所有資料的來源類別不同。

這類來源才能提供的該組織四項資訊:

  • 公開頁面是唯一可存取的介面。「Data API 拒絕了未經驗證的請求,因此我們擷取公開頁面並儲存確切位元組。」此快照以 SHA-256 固定雜湊值(6f19f8f0…,收錄於隨附的 manifest),擷取日期為 2026-07-06。任何想取得可重現 AA 數字的人,都得擷取 HTML,而非呼叫 API。
  • **排行榜規模遠超本維基過去接觸的排行榜。**共 548 種模型組態,附有各基準測試的分數、發布日期與供應商中繼資料,涵蓋至少十四項基準測試;同一個基礎模型可能以數種推理與運算量設定出現,因此「組態」與「模型」是不同的計數(421 種組態縮減為 89 個具備完整十二項基準測試組合的不同基礎模型)。
  • **AA 在同一套 harness 下自行執行所有基準測試,並建構其中三項。**研究中的十二項基準測試全由營運方在單一 harness 下執行,而 AA-Omniscience、AA-LCR 與 Terminal-Bench Hard 由 AA 自行建構或選取子集。這項特性使其資料網格獨具價值(完全沒有廠商自我報告,正是公開分數矩陣常見的混淆因素——見 Benchmark Score Redundancy),同時也限制了它(營運方特有的題目建構,會與人們在其上測量的結構混淆)。
  • **實際上,其 Intelligence Index 是能力層級標記。**在因子空間中對 409 個模型進行分群,得到兩個依發布日期區分的群集:302 個較舊模型的 Intelligence Index 平均值為 13.7(發布日期中位數為 2025-09),107 個較新模型則為 37.1(發布日期中位數為 2026-03)。

稽核對排行榜結構的結論,一句話說完:AA 的十二項基準測試表現出近乎單一維度的組合(非對角線 Spearman ρ 平均值 = 0.79,保留一個因子,解釋 74.5% 的共同變異),其主軸與發布日期的關聯為 R² = 0.505——因此,AA 對相隔數月發布模型所做的排名,很大程度上就是依發布日期排序。完整分析見 Economic Benchmark Construct Validity。

本語料接觸過的排行榜#

排行榜出現處備註
Intelligence IndexDeep Research Agents用作能力排序;代理程式表現明確呈現非單調性
GDPval-AA(Elo,v1 與 v2)GDPval Benchmark, Claude Opus 4.8, Kimi (Moonshot AI), The Open-Weight Frontier Gap以 OpenAI 的 GDPval 為基礎建立的 Elo 排行榜;v1 與 v2 無法比較
AA-Briefcase(Elo)Kimi (Moonshot AI), The Open-Weight Frontier Gap長期任務代理程式 Elo 排行榜
Conversational DynamicsInteractivity Benchmarks, GPT-Live接近飽和:三代 OpenAI 模型依序為 95.3 → 95.7 → 97.3
Speech to Speech IndexInteractivity Benchmarks, Gemini 3.8 Live本語料首個跨廠商語音指數(2026-09)
Agentic Performance (τ-Voice)Interactivity Benchmarks, Gemini 3.8 LiveAA 執行 τ-Voice 任務系列的結果
每小時輸入音訊成本Interactivity Benchmarks, Cost-per-Task Over Cost-per-Token在工作負載上實測的成本,不是價目表——一種少見而實用的呈現方式

目前記錄的兩項風險#

**版本變更會在不知不覺中破壞可比性。**AA 重新評分了 GDPval-AA,而本維基曾將 Opus 4.8 系統卡中的 v1 數字,與後續系統卡中的 v2 數字並列引用,彷彿它們屬於同一個排行榜。完整分析見 GDPval Benchmark。一般而言,第三方維護的衍生排行榜可能在名稱不變的情況下修訂;引用它的廠商標示的是自家發布日期,而不是排行榜的日期。

第三方署名,第一方方法。Google 的 Gemini 3.8 Live 發布文章中的五張圖表,標示的發布者分別是 Artificial Analysis(三張)、Sierra(一張)與 ServiceNow(一張),但五張圖表都附上相同的方法註腳——連到 deepmind.google/models/evals-methodology/gemini-3-8-live,也就是廠商自己的頁面,而非 AA 的頁面。分數很可能來自 AA;但讀者被引導查閱的測試流程則屬於 Google。沒有任何資訊能確認由哪一方執行哪些模型,以及採用何種運算量設定;圖表中的運算量標籤(自家模型為 High,最接近的競爭者為 Medium)也沒有任何說明。

為何這對本維基很重要#

AA 在本語料中扮演其他評估機構都沒有的角色:它是競爭廠商都會引用的中立計分板,也是兩家實驗室的數字在此對上的唯一途徑。最明確的例子是 Sierra 的 τ³-Banking 排行榜:OpenAI 自行發布的系統卡與 Google 的競品圖表,都列出 gpt-live-1 + Astra 32.0%、gpt-realtime-2 10.3%——兩個立場相互對立的參與方列出完全相同的數據,是本語料首度出現逐位數字一致的跨廠商佐證(Interactivity Benchmarks)。同一節也記錄了相反案例:兩家廠商以不同名稱呈報相同組態,分別得到 67.9% 與 86.2%;因此,只有在排行榜確實共用時,共用排行榜才能提供佐證。

相關連結#

資料來源#

  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google,2026-09-15(vendor-claim):三張以 AA 為標題的圖表(Speech to Speech Index、Agentic Performance τ-Voice、Cost per Hour of Input Audio),由廠商重現,並附上廠商的方法註腳
  • Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI,2026-09-10(vendor-claim):AA Conversational Dynamics 卡片
  • One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu(Oxford Internet Institute,單一作者,未經同儕審查的預印本),arXiv 2608.29420,2026-08-29,empirical:固定雜湊值的 2026-07-06 排行榜快照、遭拒的 Data API 請求、548 種組態的規模、單一 harness 與營運方建構基準測試的事實(附錄 F),以及 Intelligence Index 的層級分群(附錄 K)。這不是 AA 出版物,而是學術第三方閱讀 AA 公開頁面的研究。
  • raw/ 中沒有任何 Artificial Analysis 出版物。上方的排行榜說明是根據重現排行榜的廠商文件,以及已整理這些排行榜的維基頁面彙整而成;對本維基而言,此處列出的每個排行榜,其組成、加權方式與執行流程都未知。
§ end
Cited by 10
  • Economic Benchmark Construct Validity×3

    Artificial Analysis — the operator whose board is the dataset here, and the first source in this…

  • Cost-per-Task Over Cost-per-Token×2

    And five days later a third party measures the unit on a workload (Google, 2026-09-15). Google's…

  • GDPval Benchmark×2

    The caveat that bites this page specifically. The GDPval column in that snapshot is the Artificial…

  • Gemini 3.8 Live×2

    The launch is argued almost entirely on third-party boards rather than internal evals — Artificial…

  • Google DeepMind×2

    The disclosure posture inverts again, and this time in the lab's favour on one axis and against it…

  • Interactivity Benchmarks×2

    Artificial Analysis — the evaluator behind the Conversational Dynamics card, the Speech to Speech…

  • Benchmark Score Redundancy

    Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author preprint) runs the…

  • Compute-Controlled Benchmarking

    Everything on this page treats compute as the missing control. Zhu (Oxford Internet Institute,…

  • GPT-Live

    Artificial Analysis — the evaluator behind its Conversational Dynamics card and behind three of the…

  • Entities — People, Orgs, Tools & Projects

    Artificial Analysis — Entity. The third-party evaluator this corpus quotes most and has almost…

Related articles
  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Measuring Beyond Accuracy Saturation

    Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…