H
Howardism
Plate IIEvals & Benchmarks中文HOWARDISM

Headroom-Closed Index (HCI)

A cross-benchmark normalization that reports how much of a benchmark's *remaining* headroom a model has closed — 0 = the 90th-percentile frontier in the benchmark's entry year, 100 = a perfect score — aggregated into per-domain annual trajectories over 393 model-benchmark observations. Its finding is that progress is uneven in both level and shape: by 2026 advanced mathematics and graduate science sit at 86.4 and 85.8 while tool agents sit at 39.9 and software engineering at 52.6, and the closing rates move in opposite directions (mathematics accelerating 32.8→53.6, multimodal collapsing 59.7→2.5). Its weakness is that the normalizer is path-dependent on when a benchmark entered the dataset, and the underlying score table is not published

Article metadata
Publication details
Published:September 18, 2026
Filed:Concept
Domain:Evals & Benchmarks
Reading:19 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Headroom-Closed Index (HCI)

Sources#

Summary#

An instrument for the question a leaderboard cannot answer: not "how well does this model score" but "how much of what was left to gain has been taken." §2.1 of The Last AI Built by Humans (arXiv 2609.11873, 2026-09-10) builds it over 393 eligible model–benchmark observations across ten capability domains for models released between 2023 and September 2026, and it is the one section of that 79-page survey that is a measurement rather than an argument.

The construction, in three steps:

1. Protocol-link families. Two results belong to the same family "only when the benchmark version and evaluation harness remain stable, or when evaluations of overlapping models provide a defensible bridge between protocols." Results from incompatible versions or harnesses stay in the audit record but are excluded from the trajectories. This is strict: the rule admitted 17 of the 33 results considered in the latest-model audit, and the 16 rejected — from Terminal-Bench, DeepSWE, CyberGym, ExploitBench, AutomationBench and BrowseComp — were "retained for reference because their benchmark versions or evaluation settings could not be linked to the plotted families."

2. A provenance-weighted consensus score where several sources report the same model under the same family, variant and mode. Base weights: benchmark-owner tables 3, independent common-harness evaluations 2.5, combined benchmark-or-model reports 2, model-author tables 1 — then first-party values are multiplied by a further 0.75. The stated purpose is "a practical sensitivity adjustment, limiting the influence of self-reported results." It is the only explicit numeric discount for vendor self-reporting in this corpus (Compute-Controlled Benchmarking argues for disclosure; this one prices it).

3. The index itself, normalizing against what was left rather than against zero:

H = 100 × (s̄ − F<sub>b,0</sub>) / (100 − F<sub>b,0</sub>), where F<sub>b,0</sub> is the 90th-percentile model score in the first year the benchmark enters the dataset.

So H = 0 is the entry-year frontier and H = 100 is a perfect score. Domain trajectories then aggregate per-benchmark 90th-percentile HCI frontiers Q<sub>b,y</sub> with a √n weight on the number of distinct models contributing to each benchmark-year frontier — "the square-root weight allows better-covered families to contribute more without letting the largest table dominate the domain value."

What it finds#

Three observations, all quoted from prose rather than from any table.

Progress differs in level and in shape#

2026 HCI by domain:

Domain2026 HCI
Cybersecurity agents91.9 (not comparable across years — see caveat)
Advanced mathematics86.4
Graduate-level science85.8
Broad knowledge77.2
Legal reasoning64.5
Multimodal reasoning62.2
Frontier academic breadth60.4
Search and terminal agents56.8
Software engineering52.6
Tool agents39.9

The annual increments ΔT are where the instrument earns its keep, because they run in opposite directions over the same period:

  • Broad knowledge rises 32.8, then 26.9, then 17.6 — steady but decelerating headroom closure.
  • Legal reasoning decays hard: +48.2 in 2024, +11.4 in 2025, +4.9 in 2026.
  • Advanced mathematics accelerates: +32.8 in 2025 → +53.6 in 2026.
  • Multimodal reasoning collapses: +59.7 in 2025 → +2.5 in 2026.
  • Tool agents jump from 8.2 to 39.9 in 2026 and remain the lowest trajectory anyway.
  • Software engineering gains 40.8 in 2025 and 11.9 in 2026.

The paper's own conclusion is the one that connects this page to Aggregate Cancellation: "A single aggregate benchmark would conceal these differences in both level and trajectory shape." Two domains moving +53.6 and +2.5 in the same year average to something that describes neither, and an index built to be comparable across benchmarks is exactly the stratification that makes the divergence visible.

Interactive capabilities retain the larger gaps#

Measured against graduate-level science at 85.8, normalized headroom closure is lower by 33.2 points for software engineering, 29.1 for search and terminal agents, and 45.9 for tool agents. The leading cybersecurity-agent trajectory at 91.9 sits 52.0 points above tool agents.

The survey's reading: "gains in bounded or readily verified environments have not transferred uniformly to long, stateful workflows," because those tasks "require planning, environment-state tracking, tool selection, result interpretation, and revision of subsequent actions" and "errors propagate across the trajectory, so data collection and evaluation must cover complete interactions." This is the time-horizon result arriving from a different instrument — METR measures the length of task an agent completes reliably; HCI measures how much of the scoring headroom has closed in the domains where length binds — and the two agree on which side of the line the residual capability sits.

One caveat the paper raises against its own headline. The cybersecurity trajectory is drawn dashed because "the later Cybench observations use changed task subsets or pass@1 aggregation," so 91.9 is not comparable across years. That is the single largest number in the table and the paper marks it unusable, which is worth recording as a point in its favour.

The extrapolation is illustrative, and labelled as such#

Observation 3 projects a post-2026 region under R<sub>d</sub> = 100 − 0.22(100 − T<sub>d,2026</sub>), i.e. every domain closes 78% of its remaining gap. Cybersecurity 91.9 → 98.2; software engineering 52.6 → 89.6; search and terminal 56.8 → 90.5; tool agents 39.9 → 86.8.

The 0.22 is a chosen constant with no derivation anywhere in the paper. Its only function is to make "domains with more unclosed headroom receive a larger extension" arithmetically true, which it is by construction. The paper is honest about this — "the endpoints illustrate the hypothesis" — and the hypothesis is the survey's motivation rather than a result: that persistent, validated self-improvement would preferentially help the domains where verification-heavy interactive workflows currently fail (RSI Autonomy Levels (B0–L5)'s L3–L4 mechanisms). Treat the curve as a diagram of an argument, not as a forecast.

What Figure 3 adds that the prose does not#

Read as an image under the compile-time two-pass rule. Two things live only in the figure:

The benchmark constituents of each domain, which the prose never names: general knowledge = MMLU-Pro · LiveBench; graduate science = GPQA Diamond; academic breadth = Humanity's Last Exam; mathematics = FrontierMath v2 Tiers 1–3; multimodal = MMMU-Pro; legal = Vals LegalBench; cybersecurity = Cybench unguided / pass@1; software engineering = LiveCodeBench · SWE-bench; tool use = BFCL v3 · τ / τ²-bench; search/terminal = BrowseComp · Terminal-Bench. This matters for reading the table above: "software engineering" is LiveCodeBench and SWE-bench, not agentic repository work at large, and "search and terminal agents" includes Terminal-Bench — the same benchmark whose own §5.2 critique (on Agent-Authored Harness Optimization) argues frontier agents are near its ceiling for reasons unrelated to capability.

A mapping of the ten domains onto autonomy bands, which is the figure's legend structure and appears nowhere in the text: Knowledge and scientific evaluation · L1–L2; Specialized reasoning · L1–L2; Software environments · L2–L3; Tool-mediated workflows · L3–L4. That is the bridge between this section and the rest of the survey — the domains with the least headroom closed are exactly the ones the paper places at the highest required autonomy level.

The x-axis is measured model release date, not year buckets, with individual models labelled (GPT-4o, Claude 3.5 Sonnet, o1, Gemini 2.5 Pro, Kimi K2, GLM-4.5, Opus 4.5, Sonnet 5, Fable 5, Kimi K3, Opus 5, GPT-5.6 Terra, GPT-5.6 Sol, Fable 5.1, GLM-5.3, GPT-6 Astra). Per-model HCI values are not legible from the plot and none is quoted here.

What to distrust#

Four limits, none of them fatal and all of them unaddressed in the paper.

The normalizer is path-dependent on entry year. F<sub>b,0</sub> is the 90th-percentile score in the benchmark's first year in the dataset, so a benchmark introduced when models were already strong starts with a high floor and compresses every subsequent gain into a smaller range, while one introduced early inflates them. Two domains can therefore differ in HCI because their benchmarks were published at different points in the capability curve, not because the capabilities diverged. Frontier academic breadth (Humanity's Last Exam, designed in 2025 to be near-unsolvable) and broad knowledge (MMLU-Pro, entering when models already scored well) sit on opposite sides of this and are compared directly in Observation 1. No sensitivity analysis on the choice of the 90th percentile, on √n weighting, or on entry-year anchoring appears anywhere.

The dataset is not published. 393 observations, their family memberships, the audit records, and the per-benchmark F<sub>b,0</sub> values are all referenced and none is included. The 17-of-33 admission statistic is the only visible evidence of how aggressive the protocol-link filter is, and it is reported for a single audit.

No uncertainty of any kind. No error bars, no confidence intervals, no run-to-run variance on any trajectory, and the domain trajectories are built from 90th-percentile order statistics over small per-year model sets — where a single strong release moves the frontier by construction. A 2.5-point increment (multimodal, 2026) and a 53.6-point one (mathematics, 2026) are reported with the same apparent precision.

Provenance weighting reduces but does not remove vendor influence. Model-author tables still carry weight 1 × 0.75 rather than 0, so a capability with only first-party reporting still moves its domain trajectory. Given how little of the disclosure the grid actually carries — effort tiers, harness identity, sampling budget — two "same benchmark, same model, same mode" results can differ by more than the weighting spread.

Connections#

  • Measuring Beyond Accuracy Saturation — the sibling response to saturation and the complementary one. That page re-instruments a single saturated benchmark along new axes (reliability, cost, scaffold contribution); this one keeps accuracy and re-scales it across benchmarks so a saturating domain and a stalled one can be compared on the same ruler. Both reject retire-and-replace; they disagree about whether the fix is more axes or a better normalizer
  • Aggregate Cancellation — the failure HCI's per-domain stratification is built to prevent, with the survey's own statement of it ("a single aggregate benchmark would conceal these differences in both level and trajectory shape") and the sharpest instance in the data: +53.6 and +2.5 in the same year
  • Compute-Controlled Benchmarking — the same critique of the single-number grid, attacked from the other side. That page argues the missing dimension is the compute budget; this one supplies a provenance discount (first-party × 0.75) and a headroom normalizer while leaving the compute axis entirely unaddressed — every HCI value pools scores taken at unknown and unequal test-time budgets
  • Benchmark Score Redundancy — the other cross-benchmark score-matrix result, and the one that should worry this page: if a public model × benchmark score matrix is effectively rank-2, then ten "capability domains" may be fewer than ten degrees of freedom, and the divergence HCI reports would be a property of the normalizer rather than of the capabilities. Nobody has run BenchPress-style completion against the HCI trajectory set
  • Economic Benchmark Construct Validity — the same worry resolved rather than repeated, and the sharper objection to this page's normalizer. Zhu's single-harness snapshot finds one factor holding 74.5% of common variance across twelve benchmarks, which sounds like the counterexample this page fears until the two quantities are separated: HCI measures normalized level and rate per domain over time, that grid measures cross-model covariance at one instant, and both are true at once — benchmarks can sit at wildly different absolute levels (mathematics 86.4, tool agents 39.9) while models rank near-identically on all of them, which is exactly what a ρ = 0.79 positive manifold looks like. The real hit is on the anchor: the leading capability axis tracks release date at R² = 0.505, and HCI's normalizer is anchored to a benchmark's entry year, so the instrument's baseline is set by the single largest confound in cross-benchmark comparison
  • Task Time-Horizon Scaling — the independent trendline that agrees on the conclusion: the residual capability gap sits in long, stateful, interactive work. METR measures it in task duration, HCI in unclosed normalized headroom, and both put tool use and agentic workflows furthest from the ceiling
  • Recursive Self-Improvement — the argument this section exists to motivate: Observation 3's illustrative extension is the claim that validated self-improvement would pay off most where headroom is least closed, which is the survey's version of that page's futures 2 and 3 with a per-domain breakdown attached
  • Item Response Theory for LLM Benchmarks — the third rescaling of a saturating accuracy number, and the one whose path-dependence is a different variable. HCI normalizes against a benchmark's entry-year 90th-percentile frontier, so its scale is path-dependent on when the benchmark was published; IRT ability normalizes against per-item difficulty and discrimination, so its scale is path-dependent on which model population was calibrated against — 3,467–4,680 Open LLM Leaderboard fine-tunes, in ATLAS's case. Both replace an absolute score with a relative one and both inherit the arbitrariness of the reference class; the difference is that IRT publishes a standard error alongside the estimate and HCI publishes neither an interval nor its underlying score table
  • RSI Autonomy Levels (B0–L5) — what the survey builds this section for: Figure 3's legend maps the ten domains onto L1–L4 bands, and Observation 3's illustrative extension is the argument that recursive self-improvement would pay off most where headroom is least closed
  • Benchmark Convergent and Discriminant Validity — the test of whether capability labels mark separate axes, run on a different set of domains. Across 48 benchmarks, reasoning, knowledge and comprehension labels are not separable by cross-model covariance (within − between ρ = −0.00), and only summarization stands apart. That bounds, without settling, whether HCI's ten domains are ten trajectories, since none of HCI's high-divergence domains is in that grid

Open Questions#

  • Does HCI's ordering survive a different entry-year anchor? The index is defined against the 90th-percentile score in a benchmark's first year in the dataset, so re-anchoring (fixed calendar year, first-release score, or a random-baseline floor) is a cheap falsification test — if the ranking of the ten domains is stable across anchors the instrument is sound, and if legal reasoning and academic breadth swap places it is measuring publication timing. Needs the unpublished score table or a reconstruction.
  • Are the ten capability domains independent enough to be trajectories? Benchmark Score Redundancy reports an 84-model × 133-benchmark public matrix as effectively rank-2, which would make most of the ten domains predictable from a handful of probes — and would mean the divergence between mathematics (+53.6) and multimodal (+2.5) in one year is either the strongest counterexample to the low-rank result in the corpus or an artifact of HCI's normalization. Not answerable from the wiki alone: settling it needs the unpublished 393-observation score table, or a reconstruction of it from public leaderboards. Partially answered 2026-09-22 by Zhu (empirical, single-author preprint) — on the ranking half, and it resolves the either/or this bullet poses into a both. On a hash-pinned single-operator snapshot of twelve benchmarks spanning economic, academic, scientific-coding and long-context blocks, one factor holds 74.5% of common variance, parallel analysis retains exactly one, and hierarchical clustering by correlation distance puts ten of the twelve in a single block. So the domains are not independent as orderings of models: knowing where a model sits on one block places it on the others. That is not a counterexample to this page's divergence, because the two measure different objects — cross-model covariance at one instant against per-domain level and rate over years — and a benchmark at 86.4 closed headroom and one at 39.9 can still order the same models identically. What it does settle is that "ten trajectories" buys ten levels, not ten degrees of freedom in model comparison. And it names the mechanism this page should worry about more than rank. The leading axis tracks release date at logistic R² = 0.505, and removing the date trend costs it 14.9 points at the configuration level (24.1 on one row per base model) — while HCI normalises against the 90th-percentile score in a benchmark's entry year, which is a release-date anchor by construction. The first bullet above (re-anchoring as a falsification test) is therefore the more urgent of the two, and Zhu supplies the method for it: regress every benchmark on release date and re-fit on the residuals. Still open as posed, on two counts — Zhu's twelve benchmarks are not HCI's ten domains (no multimodal, no legal reasoning, no cybersecurity), and nobody has run either test against the unpublished 393-observation table. Partially answered (2026-09-25) by Desai et al. (empirical, COLM 2026), on whether a capability label guarantees a separate axis. It does not. On a 53-model × 48-benchmark single-pipeline grid, reasoning, knowledge and comprehension benchmarks correlate as strongly across labels as within them (within − between ρ = −0.00, robust to a model-era split), and only summarization stands apart. So HCI's ten domain labels cannot be assumed to be ten degrees of freedom. What that grid does not contain is any of HCI's high-divergence domains: no advanced mathematics beyond MATH/GSM8K, no tool agents, no multimodal. The +53.6 against +2.5 divergence is therefore untested, and the question stays open for those domains.

Sources#

  • What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks — Desai et al., arXiv 2609.08812, 2026-09-08, COLM 2026, empirical. Cited here only in the open-question annotation and the connection above: §4.2 and Table D.4 (within − between capability ρ −0.00, era-split −0.01 / +0.00; reconciled exact against pdftotext -layout p. 43). Full treatment on Benchmark Convergent and Discriminant Validity

  • One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), arXiv 2608.29420, 2026-08-29, 25pp, empirical. Cited here only in the open-question annotation and the connection above: §5.1 (one retained factor at 74.5% of common variance; first PC 79.4% of total variance), §5.2 (release-date logistic R² = 0.505; the 74.5% → 59.6% date adjustment and its 24.1-point deduplicated counterpart), and Appendix K (ten of twelve benchmarks in one correlation-distance cluster). No HCI domain is measured there and no HCI number is revised by it. Full treatment on Economic Benchmark Construct Validity

  • The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; Shanghai Jiao Tong University, Theseus Labs, Tsinghua, ByteDance, ModelBest, Xiaohongshu, Shanghai AI Lab, Humanlaya, Agent-Native Research Lab, Frontis.AI), The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement, arXiv 2609.11873, 2026-09-10, 79pp. §2.1 only — the protocol-link family rule and its 17-of-33 admission statistic, the provenance weights (3 / 2.5 / 2 / 1, first-party × 0.75), equation 2 defining HCI, the √n-weighted domain trajectory, and Observations 1–3 with all quoted values. The document-level tier is practitioner-opinion (the bulk is an argued roadmap — see RSI Autonomy Levels (B0–L5)), but this section is a quantitative secondary re-analysis of published benchmark scores and is weighted as such: the numbers below are derived measurements, not claims. Every figure on this page is taken from prose, not from a table — no table in this document's §2 was audited against the page. Figure 3 was read as an image under the compile-time two-pass rule and is the sole source for the per-domain benchmark constituents and the L1–L4 band mapping recorded above. Limits, restated from the page body: the 393-observation dataset and per-benchmark entry-year frontiers are unpublished; there is no uncertainty estimate anywhere; the entry-year anchor is path-dependent on publication timing; the 0.22 constant in Observation 3's extrapolation is undeducted and illustrative; and the paper marks its own largest value (cybersecurity 91.9) as not comparable across years because of changed Cybench subsets and pass@1 aggregation

§ end
Cited by 13
Related articles