H
Howardism
Plate IIEvals & Benchmarks中文HOWARDISM

Economic Benchmark Construct Validity

Zhu's psychometric audit of a hash-pinned Artificial Analysis snapshot (421 configurations × 12 benchmarks, four economic; every hypothesis carried on the 96-model complete-case grid): one factor holds 74.5% of common variance and tracks release date at R²=0.505, so most of the leading 'capability' axis is calendar; the four economic benchmarks form no distinct factor under the pre-specified rule, yet leave-one-benchmark-out prediction with factors re-estimated in every fold beats a single mean index by a pooled ΔMSE of 0.037 [0.019, 0.055]. Verdict: economic benchmarks add incremental predictive information to a largely date-driven general factor without constituting a separate capability — and a gap between models released months apart is mostly a gap in release dates

Article metadata
Publication details
Published:September 22, 2026
Filed:Concept
Domain:Evals & Benchmarks
Reading:27 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Economic Benchmark Construct Validity

Sources#

Summary#

Louis Yiven Zhu, One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation (Oxford Internet Institute, arXiv 2608.29420, 2026-08-29, empirical, single author, unrefereed preprint). Leaderboards now carry an economic column — GDPval, the τ-Bench family, terminal/agentic coding suites — and that column informs procurement, regulation and labour forecasts. The question nobody had asked: does it measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve?

The answer is neither of the two clean options:

  • One dominant axis. In a maximum-likelihood EFA on the complete-case grid, the first factor holds 74.5% of common variance (the first principal component is 79.4% of total score variance — a different quantity, and the one usually quoted loosely). Horn's parallel analysis retains exactly one factor: the first observed eigenvalue is 9.53 against a random-data 95th percentile of 1.79, the second is 0.91 against 1.57.
  • That axis is substantially calendar. The first factor's scores track model release date at logistic R² = 0.505 (OLS 0.477) and it is the strongest-dating of the three factors. Regressing every benchmark on release date and re-fitting drops the first factor's share 74.5% → 59.6%, a 14.9-point fall.
  • No distinct economic factor under the pre-specified rule. On the date-adjusted data parallel analysis again retains a single factor, so there is no second block to separate. H3 is reported as not supported as specified.
  • But the economic column still earns its place out of sample. Under leave-one-benchmark-out (LOBO) prediction with factors re-estimated inside every fold, a k-factor representation beats a single mean-score index by a pooled ΔMSE = +0.037, 95% bootstrap CI [+0.019, +0.055], on three of four economic targets individually.

The paper's own one-line verdict: "Economic benchmarks measure neither one capability nor many. Most of what they record is a single general factor that is largely a release-date trend, and what remains is a small but real economic component."

The design, and why the sample is smaller than it looks#

The primary data is a snapshot of the Artificial Analysis model leaderboard captured 2026-07-06 directly from the public page and pinned by SHA-256 (6f19f8f0…, recorded in a shipped manifest). The Data API refused an unauthenticated request, so the author stored the exact bytes of the public page instead. Epoch AI's Notable AI Models set supplies training compute for a scale robustness check only. Both snapshots are pinned; the analysis is a single notebook that runs top-to-bottom from them with no live network call.

Benchmarks are treated as items and model configurations as respondents — the psychometric inversion of the usual leaderboard reading.

A coverage rule fixed in advance (a benchmark kept at ≥60 models scored; a model kept if it carries ≥8 of the 13 sufficiently covered benchmarks) reduces 548 raw configurations to 421 configurations × 12 benchmarks. APEX-Agents is dropped for sparsity (26 models); MMMU-Pro is demoted to sensitivity-only.

The number that matters most is not 421. Because the economic benchmarks are scored on far fewer models than the academic core, the paper runs two grids, and every hypothesis is carried on the complete-case grid of n = 96 models scored on all twelve. The 409-model dense grid (nine near-universal benchmarks) serves only the corroborating clustering. Anyone quoting "421 models" for the factor structure or the ΔMSE is quoting the wrong denominator — the coverage table shows why: GDPval is scored on 117 configurations (112 retained), Terminal-Bench v2.1 on 121, τ³-Banking on 112, against 513 for GPQA Diamond and 509 for HLE.

The twelve, by the paper's taxonomy: economic — GDPval (Elo), Terminal-Bench v2.1, τ³-Banking, τ²-Bench; academic — GPQA Diamond, HLE, AA-Omniscience; scientific coding — SciCode, CritPt, Terminal-Bench Hard; long-context / instruction — AA-LCR, IFBench. Three of the twelve (AA-Omniscience, AA-LCR, Terminal-Bench Hard) are constructed or subset by the leaderboard operator itself, and all twelve are run by that one operator under a single harness — the property that makes this grid different in kind from a vendor-reported one.

Factor analysis is appropriate on this grid by the usual gates: KMO = 0.933 (threshold 0.8), Bartlett's sphericity decisively rejected. The battery is already near-collinear before any modelling: mean off-diagonal Spearman ρ = 0.79, every pair positive, with the economic benchmarks correlating with the academic core almost as strongly as with one another — a stronger positive manifold than the 0.73 Ilić & Gignac report across 591 models on twelve tests.

The release-date confound#

This is the finding with the widest reach, and it is the one that transfers past the economic question.

Prior psychometric work on model scores controls for scale — Kearns residualises every score on a fitted scaling law; Ilić & Gignac correlate the general factor with parameter count at ~0.6; Ruan et al. build a capability space from observational scaling. On a frontier snapshot the calendar takes that role instead:

SpecificationDate-adjustment drop in first-factor shareClears the pre-set 15-point threshold?
Configuration level, n = 96 (primary)74.5% → 59.6% = 14.9 pts (bootstrap CI [−5.3, +32.7])no
One row per base model, n = 89 (planned check R1)24.1 pts (share rises to 90.2% before adjustment)yes
Compute-known subsample, n = 58, date only16.5 ptsyes
Compute-known subsample, date + log-compute jointly9.3 pts—

Two things to take from that table.

First, the author reports the headline as failing. H2(ii) was pre-specified at fifteen points; the primary specification returns 14.9, and the paper declines to promote the deduplicated grid — which would pass — to primary, "without moving the line after seeing the result." Two of three specifications pass, so the existence of the confound is well supported and the size of the correction is not pinned down. The primary bootstrap interval [−5.3, +32.7] is indistinguishable both from zero and from the threshold, which is the honest reading of a 96-row grid.

Second, adding compute makes the correction smaller, not larger. On the same 58 models, date alone removes 16.5 points and date-plus-log-compute removes 9.3 — plausibly because later models are also larger and the two predictors split their shared variance. The operational consequence for Compute-Controlled Benchmarking: on a frontier cross-section, controlling for compute is largely controlling for the calendar with a noisier instrument, and a study that controls for scale while ignoring release date has not removed the confound it thinks it has.

Date adjustment does not merely shrink F1 — it redistributes. Figure 2c's common-variance shares move F1 74 → 60, F2 22 → 35, F3 3 → 5: the academic factor absorbs most of what the general factor gives up.

The economic block: separates only under an over-extraction#

Under the dimensionality rule the paper fixed in advance, there is nothing to report — one factor, no blocks. The separation appears only when three factors are forced, which no criterion selects, and the paper labels it exploratory throughout.

Date-adjusted three-factor oblimin loadings (Table 3, reconciled cell-for-cell against pdftotext -layout; economic in bold):

BenchmarkF1 (agentic / work-realistic)F2 (academic knowledge)F3 (physics / hard reasoning)
GDPval (Elo)0.840.13−0.04
Terminal-Bench v2.11.01−0.120.03
τ³-Banking0.540.040.38
τ²-Bench0.500.31−0.13
GPQA Diamond0.010.950.04
HLE0.320.280.49
AA-Omniscience0.690.120.14
SciCode0.260.610.20
CritPt0.02−0.000.99
Terminal-Bench Hard0.820.100.10
AA-LCR0.340.67−0.12
IFBench−0.130.860.02

All four economic benchmarks clear 0.40 on F1 and exceed their cross-loadings (largest economic cross-loading 0.38, τ³-Banking on F3). Mean absolute F1 loading is 0.72 for the economic block against 0.32 for the rest.

Four reasons to hold this loosely, all of them the author's:

  • F1 is not an economic factor, it is an agentic one. Two benchmarks the taxonomy calls non-economic sit on it as hard as the economic ones do — Terminal-Bench Hard at 0.82 and AA-Omniscience at 0.69. The paper discloses that its own plan had grouped Terminal-Bench Hard with Terminal-Bench v2.1 as economic and that the assignment is "genuinely contestable." Whatever F1 measures, it is tool-using multi-step task completion, not paid work as such.
  • The factors are strongly correlated — 0.67 between F1 and F3, 0.72 between F2 and F3 — so the separation reads as a secondary refinement of one dominant axis rather than three capabilities.
  • The loadings are loosely estimated. Bootstrap intervals run [0.10, 0.78] for τ²-Bench and [0.44, 0.99] for GDPval: they exclude zero but overlap one another, so the block's internal ordering carries no information.
  • Independent clustering agrees the structure is thin. Hierarchical clustering of benchmarks by correlation distance puts ten of twelve — and three of four economic — in one block, with only τ²-Bench and IFBench branching off. Clustering models (dense grid, n = 409) finds k = 2 by silhouette under all three algorithms (0.480 k-means, 0.499 agglomerative, 0.467 GMM), and the two clusters are capability tiers on one axis: 302 older, lower models (median release 2025-09, mean Intelligence Index 13.7) against 107 newer, higher ones (median 2026-03, mean index 37.1). That is the date axis again, seen without a factor model.

The predictive test, and the rung that fails#

Task 2 converts distinctiveness into a forecast: hold out one economic benchmark entirely, predict it from representations of the remaining eleven, and see which representation wins. Factors are re-estimated inside each training fold and the held-out fold projected through training-fold Thurstone weights — a fit-on-all-data pipeline would leak the target into the weights used to predict it. Hyperparameters come from an inner five-fold grid search nested inside the outer five-fold split that supplies the reported error.

Predictor ladder, economic block (best learner per rung):

RungLearnerTrain RMSETest RMSEOut-of-fold R²
(i) timing / scale onlyrandom forest0.5210.7410.459
(ii) mean-score indexridge0.4630.4740.771
(iii) first factor onlyelastic net0.9330.9500.110
(iv) k factors (k = 3)ridge0.4100.4330.808
(v) k factors + covariatesridge0.3910.4380.800

ΔMSE = MSE(mean index) − MSE(k factors), paired bootstrap, B = 2000:

TargetΔMSE95% CI
GDPval (Elo)+0.026[+0.010, +0.044]
Terminal-Bench v2.1+0.028[+0.001, +0.056]
τ³-Banking+0.056[+0.008, +0.103]
τ²-Bench+0.038[−0.011, +0.086]
Pooled economic+0.037[+0.019, +0.055]

H4 is supported on the rule fixed in advance. The gain survives deduplication to 89 distinct base models (+0.038, [+0.020, +0.056]), so it is not an artefact of treating reasoning-effort variants of one model as independent draws, and it survives the factor count: sweeping k ∈ {2,3,4,5} gives +0.021 / +0.037 / +0.058 / +0.059 with pooled test R² rising 0.79 → 0.83, so the k = 3 the paper adopts is the conservative choice relative to the criterion's k = 4.

Rung (iii) is the result most worth carrying elsewhere, and the paper spends one sentence on it. The single factor that holds 74.5% of common variance predicts held-out economic scores terribly — RMSE 0.950, R² 0.110 — while a naive mean of the other eleven scores reaches R² 0.771. Dominance of a factor is therefore not sufficiency of that factor: the factor-score projection discards exactly the level information a crude average keeps. Any "the score matrix is effectively rank-k" claim (Benchmark Score Redundancy) should be read with this separation in view — the rank statement is about covariance geometry, and the useful predictor may not be the leading component of it.

On GDPval specifically, a k-factor ridge reaches out-of-fold R² = 0.88, and SHAP attribution on the matched gradient-boosting learner ranks F1 first, then the reasoning flag, then F3, with log-params and open-weights status contributing little. So the scale signal prior work reports reaches GDPval through the factors, not alongside them. The largest residual is interpretable rather than diagnostic: Grok 4.3 Non-reasoning beats its economic prediction by +1.14 z-units, which fits GDPval rewarding agentic behaviour the other benchmarks capture only indirectly.

Who should trust an economic leaderboard, and for what#

The paper is unusually explicit here, and the three audiences get three different answers.

  • Buyers. A single intelligence index captures most of what distinguishes models, so a leaderboard remains a sound guide to overall progress. The residual economic signal is nonetheless the part most relevant to claims about professional value, which is why reporting the economic column separately is justified — an improvement of 0.037 sitting on a baseline that already explains 77% of economic-block variance is incremental, but it survives a validation that rewards transfer and penalises memorisation.
  • Regulators and anyone reading a gap. "A leaderboard ranking models released months apart is ranking them largely on release date." The prescription is concrete: date-adjust the score, or restrict the comparison to one release window, before reading a small gap between contemporaneous models as a capability difference. Scoring a procurement or a regulatory threshold off an unadjusted cross-model gap is scoring the calendar.
  • Labour economists. Almost nothing, and the paper says so. Scope is fixed to internal validity — the correlational and predictive structure of aggregate scores on one cross-section. Item-level responses, external economic impact such as adoption or revenue, and saturation dynamics are explicitly outside it. This paper establishes that GDPval carries information a general index does not; it establishes nothing about whether GDPval predicts anything that happens in a labour market.

A two-test protocol for benchmark builders falls out of the design and is runnable on what a leaderboard already publishes (per-model scores on the new benchmark and the existing battery, plus release dates). Step one, structural: regress every benchmark on release date, fit an oblique factor model to the residuals, check whether the new benchmark's loading separates from the general factor. Step two, predictive: hold the new benchmark out, predict it from a k-factor representation of the others with factors re-estimated in every fold, check that it beats a single mean-score index with a bootstrap interval excluding zero. "A benchmark that passes both adds a construct, and one that passes neither is re-measuring the progress trend." The author proposes an operator publish both statistics beside every score it adds.

What to distrust#

  • Single author, arXiv preprint, no peer review. Acknowledgements thank one commenter. Nothing here has been refereed, and the paper's own sharpest results are the ones it reports as failures.
  • "Pre-specified" is doing real work, and the deposit is retrospective. The analysis plan is deposited at OSF (10.17605/OSF.IO/VD34J) after the fact, with a note recording the timeline and "the limits of that evidence." The discipline is visible in the paper's behaviour — two failed hypotheses reported as failed, five deviations enumerated, a threshold admitted to have been calibrated on a mistaken reading of Kearns and left unmoved anyway — but the pre-registration itself is not independently timestamped.
  • n = 96. Every headline rests on 96 complete cases (89 deduplicated base models) against a plan that targeted 180–220. Wide intervals throughout are the visible cost.
  • One operator, one date, three of twelve benchmarks built by that operator. The single-harness property is the study's great strength against vendor-reporting bias and its single greatest external-validity limit: nothing here has been replicated on a second leaderboard, and the paper's fourth limitation concedes the results "describe one leaderboard on one date."
  • The economic factor is an over-extraction, which the author names first among four limitations: it stands as a hypothesis for a confirmatory factor model on an independent snapshot, not as an established structure.
  • Four planned robustness checks were not run (FIML/MICE missingness sensitivity, the single-factor foil replication, reasoning-effort variants as a second scale axis, and the taxonomy swap test) — and the last two bear directly on the two deviations the paper concedes are contestable.
  • Which GDPval-AA board? The snapshot's GDPval column is the Artificial Analysis Elo board, which Artificial Analysis records as having been silently rescored between a v1 and a v2 that are not comparable. The paper does not name a version. For a within-snapshot covariance analysis this is harmless; for anyone trying to line these GDPval numbers up against a vendor card, it is not.

Connections#

  • Benchmark Score Redundancy — the same redundancy, measured on the grid that removes the confound that page could not rule out. BenchPress's 84 × 133 matrix is ~80% vendor-self-reported, and its own scope caveat flags that shared reporting bias might manufacture the correlation; this is a single-operator, single-harness cross-benchmark grid, and the redundancy is stronger there (mean pairwise ρ = 0.79, first PC 79.4% of total variance). Two things it adds rather than repeats: a confound nobody in the redundancy literature controls for (release date, which carries what scale carries across generations), and the rung-(iii) separation above — the dominant factor is a poor predictor even where it is overwhelmingly dominant, so low rank does not license predicting from the top component
  • Headroom-Closed Index (HCI) — the apparent contradiction in the corpus, and it resolves cleanly. HCI reports ten capability domains diverging wildly (advanced mathematics at 86.4 against tool agents at 39.9 by 2026; closing rates moving in opposite directions); this reports one factor governing 74.5% of common variance across twelve benchmarks. Both hold, because they measure different objects: HCI measures normalised level and rate per domain over time, this measures cross-model covariance at one instant. Benchmarks can sit at wildly different absolute levels while models rank near-identically on all of them — and that is exactly what a positive manifold with ρ = 0.79 looks like. The sharper point for HCI is that its entry-year normaliser is anchored to when a benchmark entered the dataset, which this paper shows is the single largest confound in a cross-benchmark comparison
  • Compute-Controlled Benchmarking — the control that page is built around, shown to be the second-best instrument on a frontier cross-section. Adding log-compute to a date adjustment makes the correction smaller (16.5 → 9.3 points on the same 58 models), because later models are also larger; the calendar carries what scale carries. Compute control remains the right answer to benchmark-maxxing (budget per evaluation is a fairness property of one run); it is not the right answer to cross-model comparability, where the disclosure this page's evidence demands is the release date
  • GDPval Benchmark — the economic benchmark this snapshot scores highest and predicts best, sitting on the agentic factor with Terminal-Bench rather than in a block of its own. GDPval loads 0.84 on F1, is predicted out-of-fold at R² = 0.88, gains +0.026 [+0.010, +0.044] from the multi-factor representation, and its SHAP ordering puts the agentic factor above scale — the quantitative version of that page's "the benchmark measures a low-context expert" finding, read from the covariance side
  • Evaluation Horizon Versus Release Cadence — the same calendar, as a different constraint. That page has cadence outrunning the evaluation of a model; this has cadence dominating the comparison of two. Both point at the same disclosure gap: release date is the most load-bearing uncontrolled variable on a leaderboard, and no board treats it as one
  • Measuring Beyond Accuracy Saturation — the adjacent diagnosis of why a headline accuracy under-uses a benchmark. That page's answer is to re-instrument a saturated benchmark along non-accuracy axes; this one's is to residualise the axis that saturation rides on. The link between them is Akhtar et al.'s finding that saturation rises with benchmark age, which is what makes release date a first-order confound in the first place
  • Machine Self-Report Psychometrics — the corpus's other exploratory factor analysis with models as respondents, on the opposite subject: 45 human questionnaires over 50 models yield a single dominant machine-native axis (the Pinocchio Axis, up to 47.1% of between-model variance), just as twelve benchmarks yield one capability axis at 74.5%. Both then argue the one axis is "one axis only in projection" and split it — and both splits are exploratory. The methodological rhyme is worth holding: models-as-respondents grids appear to be near-unidimensional wherever anyone has looked, whether the items are benchmarks or self-report inventories
  • Market-Priced AI Exposure (the AI Premium) — the mirror image, and the reason it is worth crossing domains for. Exposure instruments are mutually near-orthogonal (the market-implied skill map shares under 2% of variance with every task-based measure); the capability instruments those measures take as their input are near-collinear, one factor at 74.5% of common variance with ten of twelve benchmarks in a single cluster. So the disagreement between exposure measures cannot be inherited from a multidimensional capability frontier — it is manufactured in the mapping from capability to tasks, occupations and prices
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the taxonomy of instruments whose denominator this audits. Every exposure measure assumes a capability construct it never validates; this validates the construct and finds it is mostly a release-date trend, which makes "what AI can do" a moving target whose movement is largely calendar. What it does not supply is the thing that taxonomy most needs — no instrument here is scored against any realised economic outcome, because the paper fixes its scope to internal validity by design
  • Item Response Theory for LLM Benchmarks — the same psychometric audit one level down, with a compatible verdict and a matching confound. This page factor-analyses models × benchmarks and finds one factor at 74.5% of common variance that tracks release date at R² = 0.505; ATLAS fits a latent trait over models × items inside a single benchmark and finds that percent-correct misplaces 23–31% of models by more than ten ranks. Both conclude the headline number is a poor estimate of the construct underneath it. The structural echo is that each scale turns out to be relative to its reference class — here the release window the snapshot happens to span, there the Open LLM Leaderboard fine-tune population the item difficulties were calibrated against — and neither paper can say what the scale would look like against a different one
  • Artificial Analysis — the operator whose board is the dataset here, and the first source in this corpus to read one of its boards at source rather than through a vendor's reproduction
  • Cognitive Capability Profiling for Task Suitability — the one-capability-or-many question with the axes swapped, and a design built to dodge the calendar confound. Zhu factor-analyses models × benchmarks and discovers a dominant factor; Prunty et al. pre-specify the factors from cognitive theory (18 capabilities from Carroll and core-knowledge work, screened to 16 by inter-rater reliability, clustered to 8) and measure items against them, so dimensionality is a design choice rather than a finding. A rubric demand level does not move with release date, which is the structural answer to this page's R² = 0.505 result — but the shared intercept reintroduces a version of the same problem, since the anchoring of every capability level shifts with the catalogue's membership, and their catalogue is five cheap-tier models plus one Pro. Their headline that systems "differ more across cognitive dimensions than across model families" rests on a 5.30-wide dimension spread against a 1.12-wide system spread with no variance decomposition reported, which is weaker evidence than the parallel-analysis machinery here
  • Benchmark Convergent and Discriminant Validity — the construct-validity audit this page's economic question sits inside, run over 48 benchmarks and 53 models with the multitrait-multimethod lenses instead of a factor model. Its capability finding is this page's one-factor result seen from the label side. Reasoning, knowledge and comprehension benchmarks correlate as strongly across labels as within them (within − between ρ = −0.00), so a label on a leaderboard column is not evidence of a separate construct. Two things it adds. First, the non-discrimination holds inside each half of a split by release date (−0.01 before October 2024, +0.00 after), so the shared axis among capability benchmarks is not only the calendar confound this page measures. Second, the safety side of the grid behaves nothing like this one: same-label safety benchmarks barely converge (bias 0.20, safety detection 0.02). That paper names first-PC residualisation as its own untested next step. It has never residualised on release date, which is this page's prescription

Open Questions#

  • Does the date-driven structure replicate on a second, independently-operated leaderboard, or is it an Artificial Analysis artefact? Three of the twelve benchmarks are constructed or subset by the operator itself and all twelve are run under its single harness, so operator-specific item construction and harness choices are confounded with the structure. Falsifiable directly: run the same two tests on a second board (Open LLM Leaderboard, LMArena, HELM) that shares models but not items, and check whether the first factor's date R² lands near 0.505. Partially answered (2026-09-25) by Desai et al. (empirical, COLM 2026). Their grid is a second single-pipeline grid, independent of Artificial Analysis: 53 models × 48 HELM- and SafetyPrompts-derived benchmarks, all run by the authors under one zero-shot harness. On it, the capability benchmarks again share one axis, with within-label ρ no higher than across-label ρ (−0.00). So the structure is not an Artificial Analysis artefact. The date half is only bounded. The paper never regresses on release date, but it splits the models at October 2024 (n = 27 / 26) and the non-discrimination holds inside each half (−0.01 / +0.00). Within each half of the release range, capability benchmarks still do not separate. Each half spans roughly a year or more (Llama-2 chat models to September 2024, then October 2024 to at least claude-sonnet-4-5), so this is a coarse window rather than a single release burst. The result is consistent with the date trend being a large share of the first factor rather than all of it. The posed test, the first factor's date R² on a second board, is still unrun.
  • Is there a confirmatory economic factor, or only an over-extraction? The paper's own first limitation: the three-factor separation is exploratory, no criterion selects it, the loading bootstraps overlap each other, and F1 holds two benchmarks the taxonomy calls non-economic (Terminal-Bench Hard 0.82, AA-Omniscience 0.69). Falsifiable: fit a confirmatory factor model with the economic block pre-specified on an independent, later snapshot and report its fit against the one-factor alternative.
  • Is the calendar confound a permanent property of frontier leaderboards or a feature of the 2025–2026 release burst? Every model in this snapshot sits inside a window in which capability rose steeply and jointly; if releases desynchronise across labs, or progress within a generation outruns progress between them, the first factor's date R² should fall without any change to the benchmarks. The trigger to watch is the first snapshot in which date adjustment costs the first factor less than ten points at the configuration level.

Sources#

  • One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (University of Oxford, Oxford Internet Institute; single author), One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation, arXiv 2608.29420, 2026-08-29, 25pp, empirical — tier kept (pinned inputs, pre-specified thresholds, nested cross-validation, failed hypotheses reported as failed), with the standing caveat that it is an unrefereed single-author preprint whose analysis-plan deposit is retrospective. §3 (the 2026-07-06 hash-pinned Artificial Analysis snapshot, the coverage rule, the two grids, KMO 0.933 and ρ = 0.79); §4.1–4.3 (parallel analysis, oblimin EFA, date adjustment as rank-one removal, the four pre-specified hypotheses and thresholds); §5.1 (79.4% first PC, eigenvalues 9.53/1.79 and 0.91/1.57, 74.5% first-factor common variance); §5.2 (logistic R² 0.505 against OLS 0.477; 74.5% → 59.6%; the 14.9 / 24.1 / 16.5 / 9.3-point specification table); §5.3 and Appendix E Table 3 (the date-adjusted three-factor loadings); §5.4 and Table 1 (the predictor ladder and the ΔMSE bootstrap); §5.5 and Appendix L (dedup 90.2% and +0.038, rank 94.5%, logit 91.8%, the k ∈ {2,3,4,5} sweep, the four planned checks not run); §6 (the leaderboard-use guidance, the two-test protocol, the four limitations); Appendix F Table 4 (per-benchmark coverage and the operator-constructed three); Appendix J (five disclosed deviations, including the contested Terminal-Bench Hard assignment and the mis-calibrated H2(ii) threshold); Appendix K (clustering, k = 2, the 302/107 tiers); Appendix M (GDPval R² = 0.88, the SHAP ordering, the Grok 4.3 Non-reasoning +1.14 residual); Appendix N (provenance, the SHA-256 manifest, the refused Data API request, the retrospective OSF deposit). Parse notes: PDF-derived (docling 2.126.0, MLX layout + table stages, 25pp, 6 tables, 10 figures, confidence excellent); all five soft checks passed at ingest with canary-recall 20/20. Tables 1, 3, 4 and 5 were nonetheless reconciled cell-for-cell against pdftotext -layout and are exact — no collapse, shift, weld or split row anywhere in the document, which is unusual for this corpus. Figures 2 and 3 were read as images in the second pass; Figure 2c supplies the F2 22 → 35 redistribution, which appears in no table, and Figure 3c's SHAP bar ordering confirms the prose

  • What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks — Desai, Truong, Wallach, Chouldechova, Cooper, Garcia-Gathright, Ho, Jacobs, Koyejo, Pangakis & Wang, What AI Benchmarks Actually Measure, arXiv 2609.08812, 2026-09-08, COLM 2026, empirical. Cited here only in the connection and the open-question annotation above: §4.2 (capability labels non-discriminable) and Appendix D.5 Table D.4 (within − between capability ρ −0.00 overall, −0.01 / +0.00 across the October 2024 release split; reconciled exact against pdftotext -layout p. 43). Full treatment on Benchmark Convergent and Discriminant Validity

§ end
Cited by 15
Related articles