Sources#
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
- What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Summary#
Item Response Theory is the measurement model behind standardized testing: each item carries its own
difficulty b, discrimination a, and lower asymptote c, and each examinee carries a latent
ability θ on a scale that does not depend on which items were administered. Li, Tang, Chen, Cheng,
Metoyer, Hua & Chawla (Notre Dame, ATLAS, arXiv 2511.04689,
ICML 2026, empirical) apply it to LLM benchmarking at a scale nobody had reached before — 3PL calibration
over 3,467–4,680 HuggingFace Open LLM Leaderboard models per benchmark, on five benchmarks (WinoGrande,
TruthfulQA, HellaSwag, GSM8K, ARC) — and produce two findings that are usually reported as one and are
better kept apart:
- The scoring claim. Replacing percent-correct with whole-bank
θ̂reorders the leaderboard, and the reordering is more stable than the ordering it replaces. This uses the entire item bank; adaptive testing plays no part in it. - The administration claim. Selecting items by Fisher information and stopping when
SE(θ̂) ≤ τrecovers whole-bankθ̂from 30–89 items. This is an efficiency result about how many items you must run, and it says nothing about whetherθis the right scale.
Keeping them apart matters because the second is the one the abstract leads with ("up to 90% fewer items") and the first is the one that changes what a benchmark number means.
The scoring claim: same accuracy, different ability#
The mechanism is the one IRT was built for. A correct answer to a hard, discriminative item is strong
evidence of high ability; a correct answer to an easy one is expected and carries almost none. Percent-correct
weights them identically. So two models can be tied on accuracy and separated on θ:
- WinoGrande (Figure 2): two models at identical accuracy 0.833 receive
θ̂ = 1.2andθ̂ = 0.6. Model A's correct answers concentrate in difficulty bands D6–D9, Model B's in D1–D4. - ARC (Figure 7):
mera-mix-4x7Bat 0.713 accuracy takes whole-bank ability rank 270;LLaMAntino-3-ANITA-8B-Inst-DPO-ITAat 0.714 takes rank 2,612. A 0.001 accuracy gap, a 2,342-position ability gap. - HellaSwag (Figure 8): two models at identical accuracy 0.853 take ability ranks 347 and 3,074.
At population scale (Figure 3): accuracy and ability correlate at Spearman 0.99 on GSM8K and 0.96 on HellaSwag (Kendall 0.92 / 0.84), and yet 23% (GSM8K) and 31% (HellaSwag) of models shift by more than 10 rank positions. That pairing is the whole point — a 0.99 rank correlation over ~4,000 models still leaves a quarter of the field misplaced by more than ten slots, because rank correlation is dominated by the long capability range and the disagreements sit inside the dense middle and at the extremes.
The extremes are where accuracy fails outright. In the low-performing regime accuracy compresses into a
0.10–0.15 band while θ spans roughly −3 to −1; at the top, ceiling effects compress accuracy while
θ keeps resolving across 1.5 to 2.5. This is the accuracy-saturation problem stated as a measurement
property rather than as a benchmark-lifecycle problem — see Measuring Beyond Accuracy Saturation.
Is the reordering signal or noise?#
The obvious objection is that θ̂ simply adds estimation variance and shuffles ties. The paper answers it two
ways, and the second answer is much stronger than the first (Table 4, reconciled against pdftotext):
| Test | Raw accuracy | Ability θ̂ | Δ |
|---|---|---|---|
| Split-half rank stability, WinoGrande | 0.943 | 0.981 | +0.038 |
| …TruthfulQA | 0.981 | 0.992 | +0.011 |
| …GSM8K | 0.985 | 0.993 | +0.008 |
| …ARC | 0.968 | 0.977 | +0.008 |
| …HellaSwag | 0.994 | 0.996 | +0.002 |
| GSM8K ↔ MMLU Elementary Math | 0.504 | 0.738 | +0.234 |
| MMLU College Chem. ↔ MMLU High School Chem. | 0.439 | 0.904 | +0.465 |
Split-half (10 random item partitions, Spearman of the two half-bank rankings) improves everywhere but the
margins are small and the accuracy baseline was already ≥ 0.94 — this test mostly shows θ̂ is not worse.
The cross-benchmark rows are the load-bearing ones: between two MMLU chemistry subjects that ought to
measure nearly the same thing, raw-accuracy rankings agree at 0.439 and ability rankings at 0.904. That is
external convergent validity, not internal consistency, and it is the only evidence here that the
reordering tracks something real rather than a different arbitrary scale.
The administration claim: Fisher information and precision-based stopping#
ATLAS-the-algorithm is textbook computerized adaptive testing (CAT), with LLM-specific choices:
- Calibrate a 3PL model per benchmark, after filtering out items with response SD < 1%, mean accuracy
95%, or point-biserial
r_pb < 0.1. Saturated and non-discriminative items are simply discarded before measurement begins — the psychometric answer to a saturating benchmark is to delete the dead items, not to retire the benchmark.
- Partition and link. Fitting 3PL over a 5,600-item bank at once is
O(|I|³); instead split intoKnon-overlapping subsets of ≥ 100 items (K= 6–50), calibrate independently, then align the provisional scales with common-person linking via mean–sigma transformations. Every model answers every item, so the model population is the anchor set — an option human testing never has (no fatigue, no practice effects, no time limits). Complexity drops toO(K · maxₖ|Iₖ|³). - Administer adaptively. Start at
θ̂₀ = 0; after each response compute Fisher information for all unadministered items and sample randomly from the top 5 (randomesque selection, to avoid over-exposing one item type); updateθ̂by EAP; stop whenSE(θ̂) ≤ τwithτ ∈ {0.1, 0.2, 0.3}, floor 30 items, ceiling 500.
Measured against whole-bank θ̂ on held-out models (Tables 2 and 9, both reconciled):
| Benchmark (bank size) | Best ATLAS ability MAE | Items | Random-100 MAE | Best static baseline |
|---|---|---|---|---|
| WinoGrande (1,045) | 0.155 | 70 | 0.167 | MetaBench-P 0.152 @ 133 |
| TruthfulQA (627) | 0.064 | 48 | 0.103 | MetaBench-S 0.072 @ 136 |
| HellaSwag (5,600) | 0.157 | 41 | 0.240 | TinyBenchmarks 0.198 @ 97 |
| GSM8K (1,306) | 0.150 | 70 | 0.150 | MetaBench-S 0.096 @ 249 |
| ARC (839) | 0.084 | 89 | 0.183 | MetaBench-P/S 0.134 @ 145/100 |
The honest reading: ATLAS wins decisively on items per unit of error — it takes the lowest Information
Efficiency Score (IES = (MAE_method/MAE_random) × (Items/100)) on every benchmark, 0.193–0.655. It does
not dominate on error. On GSM8K its best ability MAE exactly ties the 100-item random baseline (0.150) and
loses to MetaBench-Secondary (0.096, at 249 items); on WinoGrande it essentially ties MetaBench-Primary. The
claim that survives is "same precision, far fewer items", not "more accurate".
In accuracy space it is weaker still. Reconstructing percent-correct from θ̂ with the p-IRT estimator
(Tables 3 and 10) keeps MAE below 5% everywhere — which is the paper's evidence that θ preserves the global
performance structure — but ATLAS is not the best reconstructor. On GSM8K it posts 0.039 against
Random-100's 0.026 and MetaBench-S's 0.020, and its IES exceeds 1 (1.055): at 70 items it is less
efficient than random sampling for the purpose of estimating accuracy. Ability estimation is what the
adaptive machinery is optimized for; accuracy reconstruction is a byproduct it does adequately.
What "up to 90%" actually means#
The abstract's "reduces the number of required items by up to 90%" is not a measured reduction. It is the
arithmetic of the 500-item ceiling against the largest bank ("the maximum constrains computational cost
and yields approximately 90% reduction in test length relative to full benchmarks"). The measured reductions
are far larger and are per-model variable, because test length is an output of the stopping rule rather
than a design parameter: 41/5,600 on HellaSwag is 99.3%, and at τ = 0.3 every benchmark lands on 30–32
items (95–99.5% of the bank unused). Two consequences worth carrying:
- The reduction is at fixed precision (
SE(θ̂) ≤ τon the latent scale), not at fixed correlation with the full-test score. The MAE-vs-whole-bank numbers are a post-hoc validation of that precision target, not the target itself. A reader who wants "90% fewer items at r ≥ 0.99 with the published leaderboard" is asking a question this design does not answer. - The banks are already filtered. The denominators above (627–5,600) are post-filter; the pre-filter item counts are never reported, so the reduction against the shipped benchmark is larger than stated and unquantified.
Runtime is the reason this is practical at all: 9.4–75.5 s per model end-to-end (Table 11), scaling with bank size. Test overlap between two models' item sequences stays at 11.3–23.7% and average item exposure under 12%, so the adaptive pool genuinely rotates — the contamination surface of a 41-item test is not 41 fixed items.
Psychometric hygiene, and where it stops#
The paper's sharpest methodological contribution is not ATLAS but the observation that prior IRT-based benchmark-reduction work never reported model fit, and that when you compute it for them, it is bad. Running M₂/RMSEA on TinyBenchmarks' and MetaBench's own released IRT code (Table 1):
- TinyBenchmarks: severe misfit on all five benchmarks — the consequence of estimating up to 15 latent
traits from only 395 models. (The printed cells read 364.24 / 371.49 / 646.82 / 506.60 / 369.89 under an
"RMSEA" column header; RMSEA is bounded by construction, so these are M₂ statistics printed in the wrong
column. Reconciled against
pdftotext— the error is the paper's, not the parse's. The qualitative verdict "Poor" is what to cite, not the numbers.) - MetaBench: RMSEA 0.0423–0.1389 — good on HellaSwag and GSM8K, marginal on ARC, poor on TruthfulQA.
- ATLAS: 0.0438–0.0690 — good or acceptable on all five.
So: a benchmark subset can predict full-benchmark accuracy well and still rest on an IRT fit that does not license the ability estimates it produces. Fit diagnostics should be a reporting requirement wherever IRT is used to reduce a benchmark.
Which assumptions were actually tested. Local independence, yes: Yen's Q₃ residual correlations (Table 8) put mean off-diagonal Q₃ at 0.002–0.011 and the 99th-percentile adjusted Q₃ at 0.144–0.257, with 0.2–2.1% of item pairs above the 0.20 threshold (TruthfulQA worst at 2.1%) — mild, plausibly template- or topic-driven, not a broad violation. Unidimensionality, no — there is no parallel analysis, scree, or bifactor comparison; the only handle on it is that a unidimensional 3PL achieves acceptable global fit, which is indirect. Monotonicity is assumed by the 3PL form. Parameter invariance is tested only indirectly, through the transfer experiments below.
The assumption nobody tests is about the persons, not the items. IRT calibration treats examinees as
exchangeable draws from a population. The Open LLM Leaderboard is thousands of community fine-tunes and
merges of a handful of base models — a population with extreme internal dependence, in which "person"
responses are anything but independent. Every difficulty and discrimination parameter here is defined
relative to that population, so b is "hard for the 2023–2024 open-weight fine-tune ecosystem", not "hard"
in any absolute sense, and a frontier-only population would re-scale the bank. No person-fit statistic is
reported either — which matters precisely because an aberrant response pattern (easy items wrong, hard items
right, or the reverse) is the signature of contamination, and Figure 8's low-ability twin is literally named
contaminated_proof_7b_v1.0.
Transfer beyond the calibration distribution#
Two holdouts on ARC (Table 5 — docling collapsed each method's three Setting rows into single cells; the raw carries a hand-uncollapsed version, re-verified cell-for-cell against page 9):
- Architecture-family holdout. All 321 Mixtral-family (MoE) models withheld from calibration. Ability MAE 0.084–0.120 → 0.091–0.109 — flat, and slightly better at the looser thresholds. Accuracy reconstruction MAE 0.032–0.034 → 0.044–0.045.
- Temporal holdout. Calibrate on 2,998 models released before 2024-05-01, test on 322 later ones (844 dropped for missing release dates). Ability MAE 0.084–0.117 → 0.126–0.162 — a ~50% degradation, the larger of the two shifts. Accuracy reconstruction again moves only to 0.044–0.045.
Temporal drift hurts more than architectural novelty, which is the result you would expect if item difficulty is population-relative: newer models are not a different kind of examinee, they are a shifted ability distribution against a bank calibrated on the old one. Item banks need recalibration on a release cadence, and nobody has measured how fast.
Two design choices that are less settled than they look#
- 3PL is not uniformly better than 2PL (Table 12). 3PL wins outright only on TruthfulQA. Elsewhere 2PL
often gets slightly lower ability MAE but needs substantially longer tests (WinoGrande
τ=0.1: 2PL 0.121 at 198 items vs 3PL 0.155 at 70), while 3PL usually reconstructs accuracy better. The paper is candid that thecparameter should not be read as "guessing" for an LLM — it may absorb answer priors, prompting conventions, decoding behavior, or memorization. - The reference estimator moves the headline. ATLAS's whole-bank reference uses WLE; switching it to EAP (Table 13) cuts the flagship HellaSwag MAE from 0.157 to 0.087 and WinoGrande from 0.155 to 0.136, while leaving ARC and TruthfulQA roughly unchanged. The conclusion ("ATLAS tracks whole-bank ability") survives either way, but any specific MAE quoted from this paper is conditional on an estimator choice buried in Appendix G.4.
What this changes about reading a leaderboard#
- A leaderboard tie is not an evidence tie. Two models at the same percent-correct can differ by thousands of positions on a scale that is more stable under item resampling and far more consistent across sibling benchmarks.
- Benchmark saturation is partly a scoring artifact. The ceiling that compresses the top of a leaderboard is a property of percent-correct, not of the item bank — and roughly a third of a bank's items can be discarded as uninformative before measurement even starts.
- Cheap evaluation and correct evaluation are separate purchases here. Adaptive selection buys the first; IRT scoring buys the second; you can take either without the other.
Connections#
- Measuring Beyond Accuracy Saturation — the same problem from the benchmark-lifecycle side. That page's prescription is to keep a saturated benchmark and add axes (reliability, cost, scaffold contribution, human uplift); this one keeps the single axis and changes the scale, which is cheaper — it needs no new instrumentation, only the response matrix the benchmark already produced — but recovers less. Between them they bracket the answer to retire-and-replace: re-instrument, or re-score
- Adaptive Stopping in Evaluation Sampling — the sibling stopping rule, on the orthogonal axis. OptStop
stops sampling epochs per item once a Bayesian credible interval is narrow enough; ATLAS stops
administering items per model once
SE(θ̂)is small enough. Both are precision-targeted rather than accuracy-targeted, both report savings against a fixed-budget baseline (57.2–97.3% of trials vs 89–99% of items), and neither has been run inside the other — an evaluation that stopped on both axes at once is the obvious unbuilt thing - Recurring Production Agent Evaluation — the same Rasch/2PL/Fisher-information machinery, applied to a different object. ATLAS fits its item bank once on a static leaderboard population; She & Lin fit theirs on temporally-ordered runs of one evolving production agent and additionally compare adaptive testing against historical caching and fixed subsets, deploying the fixed subset instead of their own best-MAE method (multidimensional-2PL adaptive) for operational reasons. Their ten calibration windows all draw from one system's own history — see the staleness open question below for what their 1-day-ties-4-week result does and doesn't say about it
- Benchmark Score Redundancy — the redundancy result one level up, and a live tension. That page carries an 84-model × 133-benchmark score matrix at effective rank 2, and CollabEval's finding that item-level matrices need ~16 components rather than 2; a unidimensional 3PL achieving RMSEA 0.044–0.069 on five item-level banks is evidence pulling the other way. The two are not strictly contradictory — CollabEval measures reconstruction components on models × prompts, IRT measures fit of a latent-trait model within one benchmark — but anyone claiming a dimensionality for benchmark items should reconcile them
- Economic Benchmark Construct Validity — the other psychometric audit of a leaderboard, with the opposite instrument and a compatible verdict. Zhu factor-analyses models × benchmarks and finds one factor at 74.5% of common variance that tracks release date at R² = 0.505; ATLAS fits a latent trait within a benchmark over models × items. Both conclude the headline number is a poor estimate of the construct, and Zhu's calendar confound is the cross-benchmark analogue of ATLAS's population-relative difficulty scale
- Machine Self-Report Psychometrics — psychometrics borrowed for models, in the other direction. That page's finding is that human questionnaires fail when pointed at models and needed a purpose-built instrument; here the human machinery (3PL, Fisher information, common-person linking, Yen's Q₃) transfers largely intact, because benchmark items are genuinely test items and models genuinely have response patterns. The transfer succeeds on items and is least examined on persons — the exchangeability assumption that a leaderboard full of merges of one base model plainly violates
- Headroom-Closed Index (HCI) — the other rescaling of a saturating accuracy number. HCI renormalizes against a
benchmark's entry-year 90th-percentile frontier and is path-dependent on when the benchmark was published;
θrenormalizes against item difficulty and is path-dependent on which model population was calibrated against. Both replace an absolute score with a relative one; they differ in what the relativity is to - Cognitive Capability Profiling for Task Suitability — the same measurement family with the difficulty parameter sourced the opposite way, and the cleanest available test of this page's population-relativity worry. ATLAS estimates 3PL difficulty
bfrom ~4,000 Open LLM Leaderboard models, sobmeans "hard for the 2023–2024 fine-tune ecosystem" and drifts ~50% in ability MAE across one calibration year. Prunty et al.'s Measurement Layouts annotate per-item demand from expert-written rubrics applied by an LLM judge (δ_jk = e^(λD_jk), λ = 1), so item difficulty never touches a response matrix and the bank cannot go stale relative to a model population. The price moves rather than disappearing: their capability levels are identified only through an intercept shared across the fitted catalogue, so adding or removing a system shifts the reference point and levels cannot be compared across catalogues, while profile shape is identified regardless (recovery r = 0.92 under every intercept treatment against level r = 0.12 with a free intercept). Neither approach has produced an absolute capability scale — only two different relativities. It also reproduces this page's reordering effect on a six-system catalogue: GPT-4o-mini is last of six on raw battery accuracy (48.7%) and fourth on inferred capability - Cross-Model Error Entanglement — the measurement of the assumption this page says nobody tests: IRT treats examinees as exchangeable, and that page's difficulty-conditioned null finds excess co-failure concentrated in exactly the intra-family fine-tune pairs that make up the Open LLM Leaderboard population the bank is calibrated on
- Benchmark Convergent and Discriminant Validity — IRT used as a validity diagnostic rather than a scorer. Desai et al. fit a 1PL with one shared trait and a 1PL with one trait per benchmark on every pair of 37 binary-item benchmarks, and read held-out ΔAUC as a discrimination test. ΔAUC near zero means one ability explains both; the reasoning–knowledge pairs sit at 0.011, and the refusal–over-refusal pairs at 0.062. The same machinery that ATLAS uses to place a model on one benchmark's scale is used there to ask whether two benchmarks share a scale at all. The two uses rest on the assumption this page flags as untested, unidimensionality within a benchmark. That paper names the problem directly (citing Jiang et al. 2026, "Can we trust item response theory for AI evaluation?") and treats an item-level non-convergence as a flag, not a verdict
Open Questions#
-
Does the ability ranking predict anything downstream that accuracy does not? The validity evidence here is entirely internal-plus-convergent — split-half stability and agreement between sibling benchmarks. Nobody has checked whether
θ̂beats percent-correct at predicting an external outcome (held-out benchmark performance, human preference, deployment success). Falsifiable directly: rank models byθ̂and by accuracy on benchmark A, and compare which predicts benchmark B — a strictly stronger test than the cross-benchmark rank correlation reported here. -
How fast does a calibrated item bank go stale? Temporal holdout is the paper's largest degradation (ability MAE 0.084–0.117 → 0.126–0.162 across a single ~1-year boundary) and it is measured at exactly one cut point, on one benchmark, with 844 of 4,164 models dropped for missing release dates. The recalibration cadence an item bank needs is unmeasured, and it determines whether IRT scoring is a standing instrument or a one-off analysis. Related evidence (2026-09-25), on a different axis of the same question, from She & Lin (
empirical). Their difficulty-stratified subset, refit on ten nested calibration windows of one production agent's own run history (287 runs down to 14, all ending the same day), shows no monotonic loss of fidelity as the window shrinks — a 1-day window (14 runs) reaches 1.41pp MAE at k=200, indistinguishable from the full 4-week window's 1.40pp. This does not answer this bullet's question: it measures how far back a calibration window must look within one system's history, not how long a calibration stays valid going forward across model-population drift, which is what ATLAS's ~1-year temporal holdout measures. The two could both be true — a difficulty ordering needing little lookback to fit, while still decaying once the fitted system changes underneath it — and the authors' own caveat (their agent was at a "relatively mature development stage" with limited day-to-day change) means even their narrower claim is untested on a fast-changing system. See Recurring Production Agent Evaluation. -
Does calibration survive a population that is not thousands of merges of a few base models? Every item parameter here is estimated on the Open LLM Leaderboard, whose "examinees" are massively non-independent, and no person-fit statistic is reported. Falsifiable: recalibrate on a deduplicated or frontier-only model set and check whether item difficulty rank order is preserved — if
bre-orders, the scale is a property of the 2023–2024 fine-tune ecosystem rather than of the items. Partially answered (2026-09-23), on whether the dependency is removable rather than on how large it is. Prunty et al. (empirical) run a working Bayesian IRT pipeline over 19,535 items in which item difficulty is never estimated from a response matrix at all — it is read off expert-written rubrics by two LLM annotators as a per-capability demand level, entering the model asδ_jk = e^(λD_jk)with λ fixed. So a difficulty scale independent of the examinee population is constructible, and the staleness mechanism in the bullet above disappears with it. What the construction shows is that the relativity moves rather than vanishing: because their pooling rule depends only onc_k − λD_jk, a free per-agent intercept leaves overall capability level unidentified (recovery r = 0.12) and they must estimate a single intercept shared across the fitted catalogue to recover it (r = 0.98) — after which levels are anchored to that catalogue's membership and the paper forbids comparing them across separately fitted catalogues. Profile shape recovers at r = 0.92 under every intercept treatment, so shape is the population-free quantity and level is not. The posed test here is untouched: nobody has recalibrated ATLAS'sbon a deduplicated or frontier-only set, and the rubric route does not measure how much ATLAS's numbers would move. See the cognitive-capability-profiling page.
Sources#
-
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent — Yining She (CMU, work done at Meta) & Lei Lin (Meta), arXiv 2609.21267, 2026-09-18, 28pp,
empirical. Cited here only for the calibration-window annotation above (§6.2, Table 12, Figure 7). Full treatment on Recurring Production Agent Evaluation -
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks — Desai et al., arXiv 2609.08812, 2026-09-08, COLM 2026,
empirical. Cited here only in the connection above: §3.2 and Appendix D.3 (the M1 shared-trait against M2 per-benchmark-trait 1PL pair, 20% held-out cells, ΔAUC with item bootstrap) and §4.2 prose (0.011 reasoning–knowledge; 0.062 refusal–over-refusal). Full treatment on Benchmark Convergent and Discriminant Validity -
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks — Peiyu Li, Xiuxiu Tang (equal contribution), Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua & Nitesh V. Chawla (University of Notre Dame), Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks, arXiv 2511.04689, 2025-10-26, ICML 2026 (PMLR 306), 24 pages,
empirical. Cited here for §3 (3PL calibration, partition + common-person linking, WLE/EAP, randomesque top-5 selection, 30/500 bounds), §4.2–4.6 and Tables 1–5, and Appendices C (filtering + Q₃), G.1 (Tables 9–10), G.2 (Table 11), G.3 (Table 12) and G.4 (Table 13); Figures 2, 3, 7 and 8 read under the image two-pass rule. Parse warning, in this paper's favour on everything cited: the raw is docling-derived (2.126.0 / docling-mlx 0.1.1,confidence_grade: excellent) and ingest reportedwarn, withtable-collapseflagging 312 cells and the ingest worker standing down before reconciling any of them. All thirteen non-MMLU tables (1–13) were re-reconciled cell-for-cell at compile time againstpdftotext -layouton (pp. 5–9, 12–14, 17–19) and match exactly; Table 5's collapse was repaired in the raw at ingest and the repair verified against page 9. Table 14 (per-MMLU-subject results, pp. 20–22) is badly shifted and split and was not reconciled — no number from it is cited on this page or anywhere in the wiki. Two further raw-parse gaps recorded at compile time: Figure 8's caption is missing entirely from the docling body (recovered frompdftotext -layout -f 24), and Table 7's caption is emitted below its table rather than above. One arithmetic discrepancy in the paper itself, unresolved: the family-holdout text says item parameters were calibrated on "the remaining 3,296 models" after withholding 321 Mixtral models, which reconciles with neither ARC's 3,747 calibration models nor its 4,164 total (the temporal split's 2,998 + 322 + 844 = 4,164 does reconcile). Treated asempirical— tier kept
Cited by 13
- Cognitive Capability Profiling for Task Suitability×5
Irt For Llm Benchmarks changes the scale; this one changes the unit of measurement — it reorganises
- Measuring Beyond Accuracy Saturation×4
Cognitive Capability Profiling — the third answer to "an aggregate score tells you little," and the…
- Recurring Production Agent Evaluation×3
Irt For Llm Benchmarks — shares the Rasch/2PL/Fisher-information machinery (ATLAS also fits an IRT…
- Benchmark Convergent and Discriminant Validity×2
Item level. A diagnostic use of item response theory. For each benchmark pair, fit a 1PL model with…
- Open Questions Backlog×2
Irt For Llm Benchmarks: Does calibration survive a population that is not thousands of merges of a…
- Adaptive Stopping in Evaluation Sampling
Irt For Llm Benchmarks — the sibling stopping rule on the orthogonal axis, and the one that stops…
- Benchmark Score Redundancy
Irt For Llm Benchmarks — the same redundancy claim at the item level, reached from psychometrics…
- Cross-Model Error Entanglement
Irt For Llm Benchmarks — where the dependence this page measures goes unmodelled: an IRT bank…
- Economic Benchmark Construct Validity
Irt For Llm Benchmarks — the same psychometric audit one level down, with a compatible verdict and…
- Headroom-Closed Index (HCI)
Irt For Llm Benchmarks — the third rescaling of a saturating accuracy number, and the one whose…
- Machine Self-Report Psychometrics
Irt For Llm Benchmarks — human psychometrics pointed at models a second time, and it transfers…
- Evals & Benchmarks
Irt For Llm Benchmarks — Score a benchmark with a 3PL IRT model instead of percent-correct and two…
- Open Questions Dashboard
Irt For Llm Benchmarks: Does the ability ranking predict anything downstream that accuracy does…
Related articles
- Benchmark Score Redundancy
Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix comp…
- Economic Benchmark Construct Validity
Zhu's psychometric audit of a hash-pinned Artificial Analysis snapshot (421 configurations × 12 benchmarks, four econom…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Cognitive Capability Profiling for Task Suitability
Prunty et al. (Cambridge CFI) put AI systems and workplace tasks in one cognitive space: rubric-annotate 19,535 benchma…
- Benchmark Convergent and Discriminant Validity
Desai, Wallach, Chouldechova, Koyejo et al. (COLM 2026) borrow the multitrait-multimethod test from social science and…
