Sources#
- Build more natural voice experiences with GPT‑Live‑1 in the API
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
What it is#
An independent AI benchmarking outfit that publishes model leaderboards and derived indices. It is the most-quoted third-party evaluator in this wiki and the one with the weakest provenance chain: as of 2026-09-21 no Artificial Analysis publication had been ingested into raw/, and none has been since — every AA number reaching this wiki through a product announcement arrives second-hand, through a vendor — a system card, a technical report, or a launch post that reproduces AA's chart. That is worth stating plainly, because a third-party board quoted by the party it ranks is not the same evidence as a third-party board read at source.
Read at source, once (2026-09-22)#
As of 2026-09-21 no Artificial Analysis publication has been ingested into (superseded 2026-09-22.) Zhu (Oxford Internet Institute, arXiv 2608.29420, raw/.empirical, single-author preprint) is the first source in this corpus whose primary dataset is an AA board, and it arrives from an academic third party rather than from a vendor reproducing a chart. It is still not an AA publication — AA has published nothing here — but the numbers are read off AA's own page rather than off a system card, which is a different provenance class from everything in the table below.
Four facts about the organisation that only a source like this could supply:
- The public page is the only accessible surface. "The Data API refused an unauthenticated request, so we captured the public page and stored the exact bytes." The snapshot is pinned by SHA-256 (
6f19f8f0…, in a shipped manifest), captured 2026-07-06. Anyone wanting reproducible AA numbers is scraping HTML, not calling an API. - The board is much bigger than the boards this wiki has met. 548 model configurations with per-benchmark scores, release dates and provider metadata, across at least fourteen benchmarks; one base model appears under several reasoning and effort settings, which is why "configurations" and "models" are different counts (421 configurations reduce to 89 distinct base models with a complete twelve-benchmark battery).
- AA runs every benchmark itself, under one harness — and builds three of them. All twelve benchmarks in the study are run by the operator under a single harness, and AA-Omniscience, AA-LCR and Terminal-Bench Hard are constructed or subset by AA itself. That is the property that makes its grid uniquely valuable (no vendor self-reporting anywhere, which is the confound that dogs public score matrices — see Benchmark Score Redundancy) and the property that limits it (operator-specific item construction is confounded with the structure anyone measures on it).
- Its Intelligence Index is a capability tier marker in practice. Clustering 409 models in factor space returns two clusters that are release-date tiers: 302 older models at mean Intelligence Index 13.7 (median release 2025-09) against 107 newer ones at 37.1 (median 2026-03).
What the audit says about the board's structure, in one line: twelve AA benchmarks behave as a near-unidimensional battery (mean off-diagonal Spearman ρ = 0.79, one retained factor at 74.5% of common variance) whose leading axis tracks release date at R² = 0.505 — so an AA ranking of models released months apart is substantially a ranking by release date. Full treatment on Economic Benchmark Construct Validity.
The boards this corpus has met#
| Board | Where it appears | Notes |
|---|---|---|
| Intelligence Index | Deep Research Agents | used as a capability ordering; agent performance is explicitly non-monotonic in it |
| GDPval-AA (Elo, v1 and v2) | GDPval Benchmark, Claude Opus 4.8, Kimi (Moonshot AI), The Open-Weight Frontier Gap | an Elo board built on OpenAI's GDPval; v1 and v2 are not comparable |
| AA-Briefcase (Elo) | Kimi (Moonshot AI), The Open-Weight Frontier Gap | long-horizon agentic Elo |
| Conversational Dynamics | Interactivity Benchmarks, GPT-Live | near-saturated: 95.3 → 95.7 → 97.3 across three OpenAI generations |
| Speech to Speech Index | Interactivity Benchmarks, Gemini 3.8 Live | the corpus's first cross-vendor voice index (2026-09) |
| Agentic Performance (τ-Voice) | Interactivity Benchmarks, Gemini 3.8 Live | AA's run of the τ-Voice task family |
| Cost per hour of input audio | Interactivity Benchmarks, Cost-per-Task Over Cost-per-Token | a measured cost on a workload, not a price list — an unusual and useful shape |
Two hazards recorded so far#
Version churn breaks comparability silently. AA rescored GDPval-AA, and the wiki had been quoting v1 figures from Opus 4.8's system card alongside v2 figures from later cards as though they were one board. Full treatment on GDPval Benchmark. The general form: a derivative board maintained by a third party can be revised under a stable name, and vendors quoting it date-stamp their own release, not the board.
Third-party attribution, first-party methodology. All five charts in Google's Gemini 3.8 Live post are titled with their publisher (Artificial Analysis on three, Sierra on one, ServiceNow on one) and all five carry the same methodology footnote — pointing at deepmind.google/models/evals-methodology/gemini-3-8-live, the vendor's page, not AA's. The scores may well be AA's; the protocol a reader is sent to is Google's. Nothing establishes which party ran which model, at which effort setting, and the effort labels on the bars (home model at High, nearest competitor at Medium) are not explained anywhere.
Why it matters to this wiki#
AA occupies a role no other evaluator in the corpus does: it is the neutral-ground scoreboard that competing vendors both quote, which is the only mechanism by which two labs' numbers have ever lined up here. The clearest instance is Sierra's τ³-Banking board, where OpenAI's self-published card and Google's competitor chart give gpt-live-1 + Astra 32.0% and gpt-realtime-2 10.3% — identical rows from two adversarial parties, which is the corpus's first digit-for-digit cross-vendor corroboration (Interactivity Benchmarks). The same section records the counter-case: two vendors reporting the same configuration at 67.9% and 86.2% under two different names, so shared boards buy corroboration only where the board is genuinely shared.
Connections#
- Economic Benchmark Construct Validity — the only external audit of an AA board in this corpus, and the first AA data read at source: a hash-pinned 2026-07-06 snapshot of 548 configurations, twelve boards treated as psychometric items, one dominant factor tracking release date
- Benchmark Score Redundancy — why AA's single-harness, no-vendor-self-reporting grid is the score matrix everyone in that literature wanted and nobody had
- Interactivity Benchmarks — where the voice-side boards are catalogued, weighted and reconciled
- GDPval Benchmark — the primary benchmark underneath the Elo board, and the v1/v2 non-comparability
- Measuring Beyond Accuracy Saturation — Elo boards are the instrument the wiki reaches for once accuracy stops separating models, and AA supplies most of them
- Compute-Controlled Benchmarking — AA's charts label a reasoning-effort setting per bar, which is more compute disclosure than most grids carry and still not a controlled comparison
- Gemini 3.8 Live, GPT-Live — the two voice products whose launches are argued on its boards
- Kimi (Moonshot AI), Claude Opus 4.8, The Open-Weight Frontier Gap — model pages that quote its Elo boards from vendor cards
Sources#
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google, 2026-09-15 (
vendor-claim): three AA-titled charts (Speech to Speech Index, Agentic Performance τ-Voice, Cost per Hour of Input Audio), vendor-reproduced, with a vendor methodology footnote - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim): the AA Conversational Dynamics card - One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), arXiv 2608.29420, 2026-08-29,
empirical: the hash-pinned 2026-07-06 leaderboard snapshot, the refused Data API request, the 548-configuration scale, the single-harness and operator-constructed-benchmark facts (Appendix F), and the Intelligence Index tier clustering (Appendix K). Not an AA publication — an academic third party reading AA's public page. - No Artificial Analysis publication is in
raw/. The board descriptions above are assembled from the vendor documents that reproduce them and from the wiki pages that already catalogue them; treat composition, weighting and run protocol for every board listed here as unknown to this wiki.
Cited by 10
- Economic Benchmark Construct Validity×3
Artificial Analysis — the operator whose board is the dataset here, and the first source in this…
- Cost-per-Task Over Cost-per-Token×2
And five days later a third party measures the unit on a workload (Google, 2026-09-15). Google's…
- GDPval Benchmark×2
The caveat that bites this page specifically. The GDPval column in that snapshot is the Artificial…
- Gemini 3.8 Live×2
The launch is argued almost entirely on third-party boards rather than internal evals — Artificial…
- Google DeepMind×2
The disclosure posture inverts again, and this time in the lab's favour on one axis and against it…
- Interactivity Benchmarks×2
Artificial Analysis — the evaluator behind the Conversational Dynamics card, the Speech to Speech…
- Benchmark Score Redundancy
Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author preprint) runs the…
- Compute-Controlled Benchmarking
Everything on this page treats compute as the missing control. Zhu (Oxford Internet Institute,…
- GPT-Live
Artificial Analysis — the evaluator behind its Conversational Dynamics card and behind three of the…
- Entities — People, Orgs, Tools & Projects
Artificial Analysis — Entity. The third-party evaluator this corpus quotes most and has almost…
Related articles
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Cost-per-Task Over Cost-per-Token
Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
