Sources#
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Summary#
An instrument for the question a leaderboard cannot answer: not "how well does this model score" but "how much of what was left to gain has been taken." §2.1 of The Last AI Built by Humans (arXiv 2609.11873, 2026-09-10) builds it over 393 eligible model–benchmark observations across ten capability domains for models released between 2023 and September 2026, and it is the one section of that 79-page survey that is a measurement rather than an argument.
The construction, in three steps:
1. Protocol-link families. Two results belong to the same family "only when the benchmark version and evaluation harness remain stable, or when evaluations of overlapping models provide a defensible bridge between protocols." Results from incompatible versions or harnesses stay in the audit record but are excluded from the trajectories. This is strict: the rule admitted 17 of the 33 results considered in the latest-model audit, and the 16 rejected — from Terminal-Bench, DeepSWE, CyberGym, ExploitBench, AutomationBench and BrowseComp — were "retained for reference because their benchmark versions or evaluation settings could not be linked to the plotted families."
2. A provenance-weighted consensus score where several sources report the same model under the same family, variant and mode. Base weights: benchmark-owner tables 3, independent common-harness evaluations 2.5, combined benchmark-or-model reports 2, model-author tables 1 — then first-party values are multiplied by a further 0.75. The stated purpose is "a practical sensitivity adjustment, limiting the influence of self-reported results." It is the only explicit numeric discount for vendor self-reporting in this corpus (Compute-Controlled Benchmarking argues for disclosure; this one prices it).
3. The index itself, normalizing against what was left rather than against zero:
H = 100 × (s̄ − F<sub>b,0</sub>) / (100 − F<sub>b,0</sub>), where F<sub>b,0</sub> is the 90th-percentile model score in the first year the benchmark enters the dataset.
So H = 0 is the entry-year frontier and H = 100 is a perfect score. Domain trajectories then aggregate per-benchmark 90th-percentile HCI frontiers Q<sub>b,y</sub> with a √n weight on the number of distinct models contributing to each benchmark-year frontier — "the square-root weight allows better-covered families to contribute more without letting the largest table dominate the domain value."
What it finds#
Three observations, all quoted from prose rather than from any table.
Progress differs in level and in shape#
2026 HCI by domain:
| Domain | 2026 HCI |
|---|---|
| Cybersecurity agents | 91.9 (not comparable across years — see caveat) |
| Advanced mathematics | 86.4 |
| Graduate-level science | 85.8 |
| Broad knowledge | 77.2 |
| Legal reasoning | 64.5 |
| Multimodal reasoning | 62.2 |
| Frontier academic breadth | 60.4 |
| Search and terminal agents | 56.8 |
| Software engineering | 52.6 |
| Tool agents | 39.9 |
The annual increments ΔT are where the instrument earns its keep, because they run in opposite directions over the same period:
- Broad knowledge rises 32.8, then 26.9, then 17.6 — steady but decelerating headroom closure.
- Legal reasoning decays hard: +48.2 in 2024, +11.4 in 2025, +4.9 in 2026.
- Advanced mathematics accelerates: +32.8 in 2025 → +53.6 in 2026.
- Multimodal reasoning collapses: +59.7 in 2025 → +2.5 in 2026.
- Tool agents jump from 8.2 to 39.9 in 2026 and remain the lowest trajectory anyway.
- Software engineering gains 40.8 in 2025 and 11.9 in 2026.
The paper's own conclusion is the one that connects this page to Aggregate Cancellation: "A single aggregate benchmark would conceal these differences in both level and trajectory shape." Two domains moving +53.6 and +2.5 in the same year average to something that describes neither, and an index built to be comparable across benchmarks is exactly the stratification that makes the divergence visible.
Interactive capabilities retain the larger gaps#
Measured against graduate-level science at 85.8, normalized headroom closure is lower by 33.2 points for software engineering, 29.1 for search and terminal agents, and 45.9 for tool agents. The leading cybersecurity-agent trajectory at 91.9 sits 52.0 points above tool agents.
The survey's reading: "gains in bounded or readily verified environments have not transferred uniformly to long, stateful workflows," because those tasks "require planning, environment-state tracking, tool selection, result interpretation, and revision of subsequent actions" and "errors propagate across the trajectory, so data collection and evaluation must cover complete interactions." This is the time-horizon result arriving from a different instrument — METR measures the length of task an agent completes reliably; HCI measures how much of the scoring headroom has closed in the domains where length binds — and the two agree on which side of the line the residual capability sits.
One caveat the paper raises against its own headline. The cybersecurity trajectory is drawn dashed because "the later Cybench observations use changed task subsets or pass@1 aggregation," so 91.9 is not comparable across years. That is the single largest number in the table and the paper marks it unusable, which is worth recording as a point in its favour.
The extrapolation is illustrative, and labelled as such#
Observation 3 projects a post-2026 region under R<sub>d</sub> = 100 − 0.22(100 − T<sub>d,2026</sub>), i.e. every domain closes 78% of its remaining gap. Cybersecurity 91.9 → 98.2; software engineering 52.6 → 89.6; search and terminal 56.8 → 90.5; tool agents 39.9 → 86.8.
The 0.22 is a chosen constant with no derivation anywhere in the paper. Its only function is to make "domains with more unclosed headroom receive a larger extension" arithmetically true, which it is by construction. The paper is honest about this — "the endpoints illustrate the hypothesis" — and the hypothesis is the survey's motivation rather than a result: that persistent, validated self-improvement would preferentially help the domains where verification-heavy interactive workflows currently fail (RSI Autonomy Levels (B0–L5)'s L3–L4 mechanisms). Treat the curve as a diagram of an argument, not as a forecast.
What Figure 3 adds that the prose does not#
Read as an image under the compile-time two-pass rule. Two things live only in the figure:
The benchmark constituents of each domain, which the prose never names: general knowledge = MMLU-Pro · LiveBench; graduate science = GPQA Diamond; academic breadth = Humanity's Last Exam; mathematics = FrontierMath v2 Tiers 1–3; multimodal = MMMU-Pro; legal = Vals LegalBench; cybersecurity = Cybench unguided / pass@1; software engineering = LiveCodeBench · SWE-bench; tool use = BFCL v3 · τ / τ²-bench; search/terminal = BrowseComp · Terminal-Bench. This matters for reading the table above: "software engineering" is LiveCodeBench and SWE-bench, not agentic repository work at large, and "search and terminal agents" includes Terminal-Bench — the same benchmark whose own §5.2 critique (on Agent-Authored Harness Optimization) argues frontier agents are near its ceiling for reasons unrelated to capability.
A mapping of the ten domains onto autonomy bands, which is the figure's legend structure and appears nowhere in the text: Knowledge and scientific evaluation · L1–L2; Specialized reasoning · L1–L2; Software environments · L2–L3; Tool-mediated workflows · L3–L4. That is the bridge between this section and the rest of the survey — the domains with the least headroom closed are exactly the ones the paper places at the highest required autonomy level.
The x-axis is measured model release date, not year buckets, with individual models labelled (GPT-4o, Claude 3.5 Sonnet, o1, Gemini 2.5 Pro, Kimi K2, GLM-4.5, Opus 4.5, Sonnet 5, Fable 5, Kimi K3, Opus 5, GPT-5.6 Terra, GPT-5.6 Sol, Fable 5.1, GLM-5.3, GPT-6 Astra). Per-model HCI values are not legible from the plot and none is quoted here.
What to distrust#
Four limits, none of them fatal and all of them unaddressed in the paper.
The normalizer is path-dependent on entry year. F<sub>b,0</sub> is the 90th-percentile score in the benchmark's first year in the dataset, so a benchmark introduced when models were already strong starts with a high floor and compresses every subsequent gain into a smaller range, while one introduced early inflates them. Two domains can therefore differ in HCI because their benchmarks were published at different points in the capability curve, not because the capabilities diverged. Frontier academic breadth (Humanity's Last Exam, designed in 2025 to be near-unsolvable) and broad knowledge (MMLU-Pro, entering when models already scored well) sit on opposite sides of this and are compared directly in Observation 1. No sensitivity analysis on the choice of the 90th percentile, on √n weighting, or on entry-year anchoring appears anywhere.
The dataset is not published. 393 observations, their family memberships, the audit records, and the per-benchmark F<sub>b,0</sub> values are all referenced and none is included. The 17-of-33 admission statistic is the only visible evidence of how aggressive the protocol-link filter is, and it is reported for a single audit.
No uncertainty of any kind. No error bars, no confidence intervals, no run-to-run variance on any trajectory, and the domain trajectories are built from 90th-percentile order statistics over small per-year model sets — where a single strong release moves the frontier by construction. A 2.5-point increment (multimodal, 2026) and a 53.6-point one (mathematics, 2026) are reported with the same apparent precision.
Provenance weighting reduces but does not remove vendor influence. Model-author tables still carry weight 1 × 0.75 rather than 0, so a capability with only first-party reporting still moves its domain trajectory. Given how little of the disclosure the grid actually carries — effort tiers, harness identity, sampling budget — two "same benchmark, same model, same mode" results can differ by more than the weighting spread.
Connections#
- Measuring Beyond Accuracy Saturation — the sibling response to saturation and the complementary one. That page re-instruments a single saturated benchmark along new axes (reliability, cost, scaffold contribution); this one keeps accuracy and re-scales it across benchmarks so a saturating domain and a stalled one can be compared on the same ruler. Both reject retire-and-replace; they disagree about whether the fix is more axes or a better normalizer
- Aggregate Cancellation — the failure HCI's per-domain stratification is built to prevent, with the survey's own statement of it ("a single aggregate benchmark would conceal these differences in both level and trajectory shape") and the sharpest instance in the data: +53.6 and +2.5 in the same year
- Compute-Controlled Benchmarking — the same critique of the single-number grid, attacked from the other side. That page argues the missing dimension is the compute budget; this one supplies a provenance discount (first-party × 0.75) and a headroom normalizer while leaving the compute axis entirely unaddressed — every HCI value pools scores taken at unknown and unequal test-time budgets
- Benchmark Score Redundancy — the other cross-benchmark score-matrix result, and the one that should worry this page: if a public model × benchmark score matrix is effectively rank-2, then ten "capability domains" may be fewer than ten degrees of freedom, and the divergence HCI reports would be a property of the normalizer rather than of the capabilities. Nobody has run BenchPress-style completion against the HCI trajectory set
- Economic Benchmark Construct Validity — the same worry resolved rather than repeated, and the sharper objection to this page's normalizer. Zhu's single-harness snapshot finds one factor holding 74.5% of common variance across twelve benchmarks, which sounds like the counterexample this page fears until the two quantities are separated: HCI measures normalized level and rate per domain over time, that grid measures cross-model covariance at one instant, and both are true at once — benchmarks can sit at wildly different absolute levels (mathematics 86.4, tool agents 39.9) while models rank near-identically on all of them, which is exactly what a ρ = 0.79 positive manifold looks like. The real hit is on the anchor: the leading capability axis tracks release date at R² = 0.505, and HCI's normalizer is anchored to a benchmark's entry year, so the instrument's baseline is set by the single largest confound in cross-benchmark comparison
- Task Time-Horizon Scaling — the independent trendline that agrees on the conclusion: the residual capability gap sits in long, stateful, interactive work. METR measures it in task duration, HCI in unclosed normalized headroom, and both put tool use and agentic workflows furthest from the ceiling
- Recursive Self-Improvement — the argument this section exists to motivate: Observation 3's illustrative extension is the claim that validated self-improvement would pay off most where headroom is least closed, which is the survey's version of that page's futures 2 and 3 with a per-domain breakdown attached
- Item Response Theory for LLM Benchmarks — the third rescaling of a saturating accuracy number, and the one whose path-dependence is a different variable. HCI normalizes against a benchmark's entry-year 90th-percentile frontier, so its scale is path-dependent on when the benchmark was published; IRT ability normalizes against per-item difficulty and discrimination, so its scale is path-dependent on which model population was calibrated against — 3,467–4,680 Open LLM Leaderboard fine-tunes, in ATLAS's case. Both replace an absolute score with a relative one and both inherit the arbitrariness of the reference class; the difference is that IRT publishes a standard error alongside the estimate and HCI publishes neither an interval nor its underlying score table
- RSI Autonomy Levels (B0–L5) — what the survey builds this section for: Figure 3's legend maps the ten domains onto L1–L4 bands, and Observation 3's illustrative extension is the argument that recursive self-improvement would pay off most where headroom is least closed
- Benchmark Convergent and Discriminant Validity — the test of whether capability labels mark separate axes, run on a different set of domains. Across 48 benchmarks, reasoning, knowledge and comprehension labels are not separable by cross-model covariance (within − between ρ = −0.00), and only summarization stands apart. That bounds, without settling, whether HCI's ten domains are ten trajectories, since none of HCI's high-divergence domains is in that grid
Open Questions#
- Does HCI's ordering survive a different entry-year anchor? The index is defined against the 90th-percentile score in a benchmark's first year in the dataset, so re-anchoring (fixed calendar year, first-release score, or a random-baseline floor) is a cheap falsification test — if the ranking of the ten domains is stable across anchors the instrument is sound, and if legal reasoning and academic breadth swap places it is measuring publication timing. Needs the unpublished score table or a reconstruction.
- Are the ten capability domains independent enough to be trajectories? Benchmark Score Redundancy reports an 84-model × 133-benchmark public matrix as effectively rank-2, which would make most of the ten domains predictable from a handful of probes — and would mean the divergence between mathematics (+53.6) and multimodal (+2.5) in one year is either the strongest counterexample to the low-rank result in the corpus or an artifact of HCI's normalization. Not answerable from the wiki alone: settling it needs the unpublished 393-observation score table, or a reconstruction of it from public leaderboards. Partially answered 2026-09-22 by Zhu (
empirical, single-author preprint) — on the ranking half, and it resolves the either/or this bullet poses into a both. On a hash-pinned single-operator snapshot of twelve benchmarks spanning economic, academic, scientific-coding and long-context blocks, one factor holds 74.5% of common variance, parallel analysis retains exactly one, and hierarchical clustering by correlation distance puts ten of the twelve in a single block. So the domains are not independent as orderings of models: knowing where a model sits on one block places it on the others. That is not a counterexample to this page's divergence, because the two measure different objects — cross-model covariance at one instant against per-domain level and rate over years — and a benchmark at 86.4 closed headroom and one at 39.9 can still order the same models identically. What it does settle is that "ten trajectories" buys ten levels, not ten degrees of freedom in model comparison. And it names the mechanism this page should worry about more than rank. The leading axis tracks release date at logistic R² = 0.505, and removing the date trend costs it 14.9 points at the configuration level (24.1 on one row per base model) — while HCI normalises against the 90th-percentile score in a benchmark's entry year, which is a release-date anchor by construction. The first bullet above (re-anchoring as a falsification test) is therefore the more urgent of the two, and Zhu supplies the method for it: regress every benchmark on release date and re-fit on the residuals. Still open as posed, on two counts — Zhu's twelve benchmarks are not HCI's ten domains (no multimodal, no legal reasoning, no cybersecurity), and nobody has run either test against the unpublished 393-observation table. Partially answered (2026-09-25) by Desai et al. (empirical, COLM 2026), on whether a capability label guarantees a separate axis. It does not. On a 53-model × 48-benchmark single-pipeline grid, reasoning, knowledge and comprehension benchmarks correlate as strongly across labels as within them (within − between ρ = −0.00, robust to a model-era split), and only summarization stands apart. So HCI's ten domain labels cannot be assumed to be ten degrees of freedom. What that grid does not contain is any of HCI's high-divergence domains: no advanced mathematics beyond MATH/GSM8K, no tool agents, no multimodal. The +53.6 against +2.5 divergence is therefore untested, and the question stays open for those domains.
Sources#
-
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks — Desai et al., arXiv 2609.08812, 2026-09-08, COLM 2026,
empirical. Cited here only in the open-question annotation and the connection above: §4.2 and Table D.4 (within − between capability ρ −0.00, era-split −0.01 / +0.00; reconciled exact againstpdftotext -layoutp. 43). Full treatment on Benchmark Convergent and Discriminant Validity -
One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), arXiv 2608.29420, 2026-08-29, 25pp,
empirical. Cited here only in the open-question annotation and the connection above: §5.1 (one retained factor at 74.5% of common variance; first PC 79.4% of total variance), §5.2 (release-date logistic R² = 0.505; the 74.5% → 59.6% date adjustment and its 24.1-point deduplicated counterpart), and Appendix K (ten of twelve benchmarks in one correlation-distance cluster). No HCI domain is measured there and no HCI number is revised by it. Full treatment on Economic Benchmark Construct Validity -
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; Shanghai Jiao Tong University, Theseus Labs, Tsinghua, ByteDance, ModelBest, Xiaohongshu, Shanghai AI Lab, Humanlaya, Agent-Native Research Lab, Frontis.AI), The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement, arXiv 2609.11873, 2026-09-10, 79pp. §2.1 only — the protocol-link family rule and its 17-of-33 admission statistic, the provenance weights (3 / 2.5 / 2 / 1, first-party × 0.75), equation 2 defining HCI, the √n-weighted domain trajectory, and Observations 1–3 with all quoted values. The document-level tier is
practitioner-opinion(the bulk is an argued roadmap — see RSI Autonomy Levels (B0–L5)), but this section is a quantitative secondary re-analysis of published benchmark scores and is weighted as such: the numbers below are derived measurements, not claims. Every figure on this page is taken from prose, not from a table — no table in this document's §2 was audited against the page. Figure 3 was read as an image under the compile-time two-pass rule and is the sole source for the per-domain benchmark constituents and the L1–L4 band mapping recorded above. Limits, restated from the page body: the 393-observation dataset and per-benchmark entry-year frontiers are unpublished; there is no uncertainty estimate anywhere; the entry-year anchor is path-dependent on publication timing; the 0.22 constant in Observation 3's extrapolation is undeducted and illustrative; and the paper marks its own largest value (cybersecurity 91.9) as not comparable across years because of changed Cybench subsets and pass@1 aggregation
Cited by 13
- RSI Autonomy Levels (B0–L5)×3
Document-level tier: practitioner-opinion, confirmed on a full read. The load-bearing contributions…
- Measuring Beyond Accuracy Saturation×2
This page's prescription and Headroom Closed Index's are both about what to do around a saturating
- Open Questions Backlog×2
Headroom Closed Index: Are the ten capability domains independent enough to be trajectories?
- Recursive Self-Improvement×2
Headroom Closed Index — the same survey's empirical half, and the motivation underneath the ladder:…
- Agent-Authored Harness Optimization
Headroom Closed Index — the measurement behind §5.2's first condition, on the benchmarks this page…
- Aggregate Cancellation
Headroom Closed Index — the corpus's largest measured instance of concealment by aggregation, and…
- Benchmark Convergent and Discriminant Validity
Headroom Closed Index — the test that page's open question asks for, run on a different set of…
- Benchmark Score Redundancy
Headroom Closed Index — the cross-benchmark score matrix put to the opposite use, and the place…
- Compute-Controlled Benchmarking
Headroom Closed Index — the only instrument in the corpus that puts a number on the…
- Economic Benchmark Construct Validity
Headroom Closed Index — the apparent contradiction in the corpus, and it resolves cleanly. HCI…
- Item Response Theory for LLM Benchmarks
Headroom Closed Index — the other rescaling of a saturating accuracy number. HCI renormalizes…
- Evals & Benchmarks
Headroom Closed Index — A cross-benchmark normalization that reports how much of a benchmark's…
- Task Time-Horizon Scaling
Headroom Closed Index — an independent instrument agreeing on where the residual capability sits.…
Related articles
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Benchmark Score Redundancy
Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix comp…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Benchmark Contamination and Decontamination
Sun, Zhan & Gales (Cambridge): per-sample distribution distances expose that aggregate-accuracy decontamination can wor…
- Cognitive Capability Profiling for Task Suitability
Prunty et al. (Cambridge CFI) put AI systems and workplace tasks in one cognitive space: rubric-annotate 19,535 benchma…
