H
Howardism
Plate IIEntities中文HOWARDISM

Artificial Analysis

The third-party evaluator this corpus quotes most and has almost never read directly — its Intelligence Index, GDPval-AA Elo board, AA-Briefcase, Conversational Dynamics, Speech to Speech Index, τ-Voice and cost-per-hour-of-input-audio charts reach the wiki only second-hand, inside vendor system cards and launch posts; two hazards recorded so far: a silent rescoring that made GDPval-AA v1 and v2 non-comparable, and a Google chart set that attributes the board to Artificial Analysis while footnoting the vendor's own methodology page. As of 2026-09-22 the provenance gap is half closed by the first third-party read at source, an Oxford audit of a hash-pinned snapshot finding one factor over 74.5% of common variance that tracks release date at R² = 0.505

Article metadata
Publication details
Published:September 21, 2026
Filed:Entity
Domain:Entities
Reading:9 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Artificial Analysis

Sources#

What it is#

An independent AI benchmarking outfit that publishes model leaderboards and derived indices. It is the most-quoted third-party evaluator in this wiki and the one with the weakest provenance chain: as of 2026-09-21 no Artificial Analysis publication had been ingested into raw/, and none has been since — every AA number reaching this wiki through a product announcement arrives second-hand, through a vendor — a system card, a technical report, or a launch post that reproduces AA's chart. That is worth stating plainly, because a third-party board quoted by the party it ranks is not the same evidence as a third-party board read at source.

Read at source, once (2026-09-22)#

As of 2026-09-21 no Artificial Analysis publication has been ingested into raw/. (superseded 2026-09-22.) Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author preprint) is the first source in this corpus whose primary dataset is an AA board, and it arrives from an academic third party rather than from a vendor reproducing a chart. It is still not an AA publication — AA has published nothing here — but the numbers are read off AA's own page rather than off a system card, which is a different provenance class from everything in the table below.

Four facts about the organisation that only a source like this could supply:

  • The public page is the only accessible surface. "The Data API refused an unauthenticated request, so we captured the public page and stored the exact bytes." The snapshot is pinned by SHA-256 (6f19f8f0…, in a shipped manifest), captured 2026-07-06. Anyone wanting reproducible AA numbers is scraping HTML, not calling an API.
  • The board is much bigger than the boards this wiki has met. 548 model configurations with per-benchmark scores, release dates and provider metadata, across at least fourteen benchmarks; one base model appears under several reasoning and effort settings, which is why "configurations" and "models" are different counts (421 configurations reduce to 89 distinct base models with a complete twelve-benchmark battery).
  • AA runs every benchmark itself, under one harness — and builds three of them. All twelve benchmarks in the study are run by the operator under a single harness, and AA-Omniscience, AA-LCR and Terminal-Bench Hard are constructed or subset by AA itself. That is the property that makes its grid uniquely valuable (no vendor self-reporting anywhere, which is the confound that dogs public score matrices — see Benchmark Score Redundancy) and the property that limits it (operator-specific item construction is confounded with the structure anyone measures on it).
  • Its Intelligence Index is a capability tier marker in practice. Clustering 409 models in factor space returns two clusters that are release-date tiers: 302 older models at mean Intelligence Index 13.7 (median release 2025-09) against 107 newer ones at 37.1 (median 2026-03).

What the audit says about the board's structure, in one line: twelve AA benchmarks behave as a near-unidimensional battery (mean off-diagonal Spearman ρ = 0.79, one retained factor at 74.5% of common variance) whose leading axis tracks release date at R² = 0.505 — so an AA ranking of models released months apart is substantially a ranking by release date. Full treatment on Economic Benchmark Construct Validity.

The boards this corpus has met#

BoardWhere it appearsNotes
Intelligence IndexDeep Research Agentsused as a capability ordering; agent performance is explicitly non-monotonic in it
GDPval-AA (Elo, v1 and v2)GDPval Benchmark, Claude Opus 4.8, Kimi (Moonshot AI), The Open-Weight Frontier Gapan Elo board built on OpenAI's GDPval; v1 and v2 are not comparable
AA-Briefcase (Elo)Kimi (Moonshot AI), The Open-Weight Frontier Gaplong-horizon agentic Elo
Conversational DynamicsInteractivity Benchmarks, GPT-Livenear-saturated: 95.3 → 95.7 → 97.3 across three OpenAI generations
Speech to Speech IndexInteractivity Benchmarks, Gemini 3.8 Livethe corpus's first cross-vendor voice index (2026-09)
Agentic Performance (τ-Voice)Interactivity Benchmarks, Gemini 3.8 LiveAA's run of the τ-Voice task family
Cost per hour of input audioInteractivity Benchmarks, Cost-per-Task Over Cost-per-Tokena measured cost on a workload, not a price list — an unusual and useful shape

Two hazards recorded so far#

Version churn breaks comparability silently. AA rescored GDPval-AA, and the wiki had been quoting v1 figures from Opus 4.8's system card alongside v2 figures from later cards as though they were one board. Full treatment on GDPval Benchmark. The general form: a derivative board maintained by a third party can be revised under a stable name, and vendors quoting it date-stamp their own release, not the board.

Third-party attribution, first-party methodology. All five charts in Google's Gemini 3.8 Live post are titled with their publisher (Artificial Analysis on three, Sierra on one, ServiceNow on one) and all five carry the same methodology footnote — pointing at deepmind.google/models/evals-methodology/gemini-3-8-live, the vendor's page, not AA's. The scores may well be AA's; the protocol a reader is sent to is Google's. Nothing establishes which party ran which model, at which effort setting, and the effort labels on the bars (home model at High, nearest competitor at Medium) are not explained anywhere.

Why it matters to this wiki#

AA occupies a role no other evaluator in the corpus does: it is the neutral-ground scoreboard that competing vendors both quote, which is the only mechanism by which two labs' numbers have ever lined up here. The clearest instance is Sierra's τ³-Banking board, where OpenAI's self-published card and Google's competitor chart give gpt-live-1 + Astra 32.0% and gpt-realtime-2 10.3% — identical rows from two adversarial parties, which is the corpus's first digit-for-digit cross-vendor corroboration (Interactivity Benchmarks). The same section records the counter-case: two vendors reporting the same configuration at 67.9% and 86.2% under two different names, so shared boards buy corroboration only where the board is genuinely shared.

Connections#

Sources#

  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google, 2026-09-15 (vendor-claim): three AA-titled charts (Speech to Speech Index, Agentic Performance τ-Voice, Cost per Hour of Input Audio), vendor-reproduced, with a vendor methodology footnote
  • Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (vendor-claim): the AA Conversational Dynamics card
  • One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), arXiv 2608.29420, 2026-08-29, empirical: the hash-pinned 2026-07-06 leaderboard snapshot, the refused Data API request, the 548-configuration scale, the single-harness and operator-constructed-benchmark facts (Appendix F), and the Intelligence Index tier clustering (Appendix K). Not an AA publication — an academic third party reading AA's public page.
  • No Artificial Analysis publication is in raw/. The board descriptions above are assembled from the vendor documents that reproduce them and from the wiki pages that already catalogue them; treat composition, weighting and run protocol for every board listed here as unknown to this wiki.
§ end
Cited by 10
  • Economic Benchmark Construct Validity×3

    Artificial Analysis — the operator whose board is the dataset here, and the first source in this…

  • Cost-per-Task Over Cost-per-Token×2

    And five days later a third party measures the unit on a workload (Google, 2026-09-15). Google's…

  • GDPval Benchmark×2

    The caveat that bites this page specifically. The GDPval column in that snapshot is the Artificial…

  • Gemini 3.8 Live×2

    The launch is argued almost entirely on third-party boards rather than internal evals — Artificial…

  • Google DeepMind×2

    The disclosure posture inverts again, and this time in the lab's favour on one axis and against it…

  • Interactivity Benchmarks×2

    Artificial Analysis — the evaluator behind the Conversational Dynamics card, the Speech to Speech…

  • Benchmark Score Redundancy

    Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author preprint) runs the…

  • Compute-Controlled Benchmarking

    Everything on this page treats compute as the missing control. Zhu (Oxford Internet Institute,…

  • GPT-Live

    Artificial Analysis — the evaluator behind its Conversational Dynamics card and behind three of the…

  • Entities — People, Orgs, Tools & Projects

    Artificial Analysis — Entity. The third-party evaluator this corpus quotes most and has almost…

Related articles
  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Measuring Beyond Accuracy Saturation

    Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…