Sources#
- A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
Question#
Two #oq/now items, answered as one synthesis because they share a failure mode:
- Cross-Model Error Entanglement: Is the BEI graph measuring entanglement or vintage? Its top 10 pairs are 2023–24 Llama and Qwen models. Difficulty is fitted from the other 16 models, so a uniformly weak pair looks co-failing everywhere. Would the intra-family signal survive a single-release-window roster, or residualisation on release date?
- Item Response Theory for LLM Benchmarks: Does the ability ranking θ̂ predict anything downstream that accuracy does not? The validity evidence is internal and convergent only. The falsification test: rank models by θ̂ and by accuracy on benchmark A, then see which ranking predicts benchmark B.
Short answer#
Both instruments read shared structure off a models × items matrix, and both condition on a quantity fitted from the same roster. BEI conditions on difficulty d_t, estimated from the other 16 models. IRT conditions on item parameters b, a, c, estimated from the calibration population. So the structure each one finds is defined relative to the roster, and three roster variables in the vault can masquerade as the construct: vintage (Zhu), family (Zheng & Yang) and capability (Li).
- Q1: a tier effect, not a family effect, and the roster cannot say which tier variable drives it. The paper's own top-10 table does not have the intra-family shape the question assumes. It is a clique of six old open-weight checkpoints, and cross-family edges sit among the intra-family ones. Vintage, weakness and base-vs-instruct status all coincide in those six models. The evidence from outside the roster points one way. Every frontier panel that conditioned on difficulty from outside itself kept its dependence, and none organised it by vendor. The head-on rerun is still unrun. Stays open; partial.
- Q2: the vault has no test of it. Everything ATLAS offers is θ̂ compared against θ̂, calibrated on one population. The vault's only out-of-sample comparison of a latent score with a raw average is Zhu's, made at the level of models × benchmarks. There the dominant factor predicts held-out benchmarks far worse than a crude mean, and three factors predict them slightly better. That sets a prior, not an answer. Stays open; partial by analogy.
Neither question can be advanced further by synthesising pages already in the vault. Both are retagged #oq/source, with the settling experiment specified below.
The shared mechanism: the conditioning variable comes from the roster#
| Instrument | What it conditions on | Where the conditioning variable comes from | Roster variable that can leak into the construct |
|---|---|---|---|
| BEI / CIG (Kuai et al.) | task difficulty d_t | failure rate among the other M − 2 models | capability and vintage: a weak pair has high fitted p_m(d) everywhere |
| 3PL θ̂ (ATLAS) | item b, a, c | ~4,000 Open LLM Leaderboard models | vintage (the temporal holdout degrades ability MAE ~50%) and family (DIF) |
| Factor scores (Zhu) | common factor | the 96-model grid | vintage: F1 tracks release date at logistic R² = 0.505 |
| Family DIF (Zheng & Yang) | 8–128-dim spectral ability | the RouterEval population | shows family is a residual that survives the adjustment |
| Kish n_eff (Kohli) | per-item difficulty | 100 human annotators (ChaosNLI entropy), not the panel | none from the panel: the conditioning is external |
| Ledoit-Wolf ρ̄ (Hossain et al.) | item difficulty | a different bank (open-weight judges) | none from the frontier roster: the conditioning is external |
| Measurement Layouts (Prunty et al.) | item demand | rubric annotation, never a response matrix | moves to the level intercept, which is catalogue-relative |
The last three rows are the constructive pattern. An instrument escapes its roster only when its conditioning variable is estimated outside that roster. Prunty's case shows the escape is partial: the dependence on the catalogue moves to the capability level instead of disappearing. That pattern organises both answers below.
Q1 — What BEI's top tier actually is#
What Table C.2 shows#
The raw's Table C.2 (A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges) lists the top 10 BEI pairs. Every one of them falls inside a set of six models: Llama-2-70b-hf, Llama-3-70B, Llama-3.1-70B, Qwen1.5-110B, Qwen1.5-72B-Chat and Qwen1.5-14B-Chat. Six models give 15 pairs, and the top 10 are 10 of those 15. The ranking, with family class:
| Rank | Pair | BEI | Class |
|---|---|---|---|
| 1 | Llama-3-70B ↔ Llama-3.1-70B | 0.0525 | intra |
| 2 | Llama-2-70b ↔ Qwen1.5-110B | 0.0398 | cross |
| 3 | Llama-3-70B ↔ Qwen1.5-110B | 0.0352 | cross |
| 4 | Qwen1.5-14B-Chat ↔ Qwen1.5-72B-Chat | 0.0333 | intra |
| 5 | Llama-2-70b ↔ Llama-3-70B | 0.0326 | intra |
| 6 | Llama-3.1-70B ↔ Qwen1.5-110B | 0.0325 | cross |
| 7 | Llama-2-70b ↔ Llama-3.1-70B | 0.0290 | intra |
| 8 | Qwen1.5-110B ↔ Qwen1.5-72B-Chat | 0.0266 | intra |
| 9 | Qwen1.5-110B ↔ Qwen1.5-14B-Chat | 0.0245 | intra |
| 10 | Llama-3.1-70B ↔ Qwen1.5-14B-Chat | 0.0211 | cross |
Four of the ten edges are cross-family, and they are not in the tail. Llama-2-70b ↔ Qwen1.5-110B is second, above every intra-family pair except the first, which is two adjacent releases of the same base weights. So inside the clique, family does not order the edges; tier does. The paper's prose ("strong intra-family entanglement (e.g., within the LLaMA family)") overstates what its own table shows. The question's premise, that there is an intra-family signal whose survival is in doubt, inherits that overstatement. (Wiki reading of the paper's table. The paper does not make this comparison.)
Three roster variables are collinear in the clique#
- Vintage. All six models were released in 2023–24. No closed-weight model appears, and neither do the roster's newest releases (GPT-5, GPT-oss-20B, Claude 4.6 Sonnet, Gemini-3.1-Pro).
- Capability. They are the roster's weak tier. This is exactly where the question's worry bites: with
d_tfitted from the other 16 models, a uniformly weak pair has a high fitted failure probability everywhere. - Post-training status. The paper's roster convention marks instruction-tuned variants with "Instruct" or "Chat" (§B.1). By that convention, four of the six are base checkpoints: Llama-2-70b-hf, Llama-3-70B, Llama-3.1-70B and Qwen1.5-110B. The one contrast in the roster that holds family and release fixed is Llama-3.1-70B base against Llama-3.1-70B-Instruct. The base appears in four of the top 10 pairs; the Instruct variant appears in none. That is suggestive only: absence from a top-10 list is not a measured value. But it is the only within-release comparison the paper supplies, and it varies post-training rather than family or date.
A single-release-window restriction, which is what the question proposes, would not separate these three on this roster. A 2024-H1 window still contains the Qwen1.5 trio and Llama-3-70B, and they are still the weak, mostly-base tier within that window. Its one cross-family pair, Llama-3-70B ↔ Qwen1.5-110B (0.0352), outranks the intra-family Qwen1.5-14B ↔ Qwen1.5-72B (0.0333). So a window restriction would keep a cross-family edge on top and still leave capability and base status unresolved.
What the other instruments add#
- Dependence survives on rosters with no vintage spread, conditioned from outside. Kohli's nine current frontier judges show φ̄ = 0.391. The majority vote falls 22.0pp short of independence even after per-item difficulty is conditioned on human annotation entropy, which comes from outside the panel. The most-correlated pairs are cross-family, and the same-family excess is only +0.047. Hossain, Yousefi & Lim's GPT-5.6-sol / Claude Opus 5 / Grok 4.5 panel correlates at 0.42 across providers against 0.40 within Gemini. Adjusting for difficulty taken from a different bank moves it only from 0.465 to 0.467. So once the conditioning variable no longer comes from the roster, dependence is still present and still does not follow vendor lines. Both are different instruments (raw verdict φ, shrunk error ρ̄) on different tasks, not BEI.
- Family-specific item behaviour exists beyond capability in the open-weight tier. Zheng & Yang find item × family residuals that survive an 8–128-dimensional ability adjustment, far richer than BEI's single
d_t. They replicate across owner-disjoint halves (median ρ =.308–.589), and they survive an owner cap of 1 and a filter removing merges, distills and LoRAs. So "family" is a real variable in open-weight response matrices, not only a proxy for weakness. It is DIF, though, not co-failure. Its sign flips across benchmarks for every family. And the paper tests no vintage variable (no release-date control appears in the raw). It shows family can matter, not that it is what BEI's top tier measures. - Capability structures judge error too. Li finds a frozen judge's task-conditioned false acceptance rising with target capability (mean Spearman +0.82 across 35 SWE-bench agents), with no lineage variable involved. This bears on Kuai's judge-bias half, not on BEI's answerer-side graph. But it is the vault's clearest evidence that capability, one of the three collinear variables above, is a live confound on the same kind of pairwise matrix.
Verdict and the settling experiment#
On the current evidence, BEI's top tier is best read as a tier effect in which vintage, weakness and base-model status are indistinguishable. It is not an intra-family effect: the table interleaves cross-family edges with intra-family ones. Three independent frontier instruments with externally sourced difficulty keep dependence without any vintage spread, which is evidence against pure vintage. Nothing in the vault separates weakness from base status.
The experiment that would settle the question has three parts, and none of them is in the vault. Rerun BEI on a roster that, within one release window, crosses (a) family with (b) base vs instruct variants of the same weights, spanning (c) a capability range. Then residualise the pair residuals on the product of the two models' accuracies before ranking. The intra-family reading survives only if same-family pairs still exceed matched cross-family pairs after that. Retagged #oq/source.
Q2 — Does θ̂ predict anything accuracy does not?#
Every piece of evidence in the vault is internal or θ̂-against-θ̂#
| Evidence | What is compared | Why it is not the posed test |
|---|---|---|
| ATLAS split-half (Table 4) | θ̂ on one half against θ̂ on the other | internal reliability; accuracy is already ≥ 0.94 |
| ATLAS cross-benchmark (0.439 → 0.904 on MMLU chemistry) | rank(θ̂_A) vs rank(θ̂_B), against rank(acc_A) vs rank(acc_B) | both sides are θ̂ calibrated on the same population, so a shared scaling artifact inflates both. The target is not independent of the scoring rule under test |
| ATLAS reordering examples (0.713 vs 0.714 → ranks 270 vs 2,612) | θ̂ rank vs accuracy rank | Zheng & Yang show any principled reweighting flips 30.9–47.1% of sub-point pairs, and a random half-subtest flips 11.4–34.7%, so a reordering is expected and is not validity evidence |
| Prunty et al. (GPT-4o-mini: last on accuracy, fourth on inferred capability) | inferred level vs accuracy | "Nothing is validated against an observed outcome" (Cognitive Capability Profiling for Task Suitability) |
| She & Lin (MD-2PL 1.03pp MAE vs Rasch 1.37pp) | IRT reconstruction of the same benchmark's full score | within-benchmark prediction; the target is accuracy on unadministered items, not an external outcome (Recurring Production Agent Evaluation) |
| Desai et al. 1PL ΔAUC | shared vs separate trait, held-out item cells | a dimensionality diagnostic, not a θ̂-vs-accuracy comparison (Benchmark Convergent and Discriminant Validity) |
The ATLAS convergence row is the closest to the test, and it carries a roster problem of its own. The raw says IRT separates models most in the tails. At the low end, accuracy compresses into a 0.10–0.15 band while θ spans about −3 to −1. The Open LLM Leaderboard population is dense in exactly that weak tail, and a Spearman over ~4,000 models puts most of its weight where most of the models are. So part of the 0.439 → 0.904 gain may be θ̂ spreading out a tail that accuracy ties at chance level. That would be a real gain in resolution, but it would sit where this roster is densest and could shrink on a frontier-only roster. Neither ATLAS nor anyone else reports the convergence gain by ability stratum. (Wiki inference from ATLAS's reported compression band and population. Untested.)
The one out-of-sample latent-vs-raw comparison, one level up#
Zhu's leave-one-benchmark-out ladder is the only place in the vault where a latent-structure score and a raw average compete to predict held-out measurements. Factors are re-estimated inside each fold:
| Predictor of a held-out economic benchmark | Out-of-fold R² |
|---|---|
| mean of the other eleven standardised scores | 0.771 |
| first factor only (74.5% of common variance) | 0.110 |
| three factors | 0.808 (ΔMSE vs mean +0.037 [+0.019, +0.055]) |
Two readings carry over to θ̂:
- A dominant latent dimension can predict worse than the crude sum it was supposed to improve on. The factor-score projection discards the level information that the mean keeps. θ̂ within one benchmark does not share that exact failure, because it is a discrimination-weighted monotone function of the responses and keeps level. So Zhu's rung (iii) is a warning about what "latent" can cost, not a prediction that θ̂ loses.
- The latent reading won only with more than one dimension, and only narrowly. The vault's dimensionality evidence points the same way within benchmarks. Zheng & Yang's held-out selection picks K = 8–128 on five item banks, two of them ATLAS's own. She & Lin's multidimensional 2PL beats Rasch at every budget. A unidimensional 3PL θ̂ may discard exactly the structure an external target needs.
This is models × benchmarks on a 96-model frontier grid, not θ̂ against accuracy within a benchmark. It says that "a latent score beats the raw score out of sample" is not free and not ruled out. It does not say which way the posed test comes out.
The test, specified so a roster artifact cannot win it#
The question's own test is right. The vault adds four constraints, each from a page that shows the failure mode:
- Score benchmark B by accuracy, not θ̂. Otherwise both sides share calibration and scaling (the ATLAS convergence row's problem).
- Exclude near-ties or model them as noise. Zheng & Yang show that sub-point orderings flip under any reweighting, so neither scoring can "win" them.
- Residualise on release date and stratify by family before comparing the two predictors. Zhu: F1 tracks date at R² 0.505. Zheng & Yang: family DIF replicates. Without this, θ̂ can beat accuracy by tracking the roster instead of the construct.
- Report the comparison by ability stratum. That shows whether any θ̂ advantage lives only in the dense low-ability tail of an Open-LLM-Leaderboard-style population.
The data exist: ATLAS's response matrices on five benchmarks, and RouterEval's matrices, which Zheng & Yang used. That makes this a re-analysis, not a new collection, but nobody in the vault has run it. Retagged #oq/source.
Tensions and contradictions found#
- Kuai et al.'s prose against their own Table C.2. The text calls BEI's leading structure intra-family ("strong intra-family entanglement (e.g., within the LLaMA family)"; Appendix C: "concentrated primarily among closely related model families"). The table has cross-family Llama↔Qwen pairs at ranks 2, 3, 6 and 10.
- Item Response Theory for LLM Benchmarks's Connections bullet on Cross-Model Error Entanglement says that page finds excess co-failure "concentrated in exactly the intra-family fine-tune pairs that make up the Open LLM Leaderboard population". Two parts of that are inaccurate. Four of BEI's top 10 are cross-family. And Kuai's roster holds official base and chat releases, not community fine-tunes. The analogy to the leaderboard population still holds at the level of "older open-weight tier". Left for the next lint pass to reword.
- Kohli against Hossain on the sign of the family gap (same-family +0.047 above cross-family, against cross 0.42 ≥ within 0.40). This is already recorded on Cross-Model Error Entanglement. Both are small and of the same order, and both support "vendor is not the decorrelating variable", so they are consistent in conclusion.
Sources#
- Cross-Model Error Entanglement: BEI/CIG construction, roster, judge-bias cautions, the Kohli and Hossain sections, and the question's annotation history.
- A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges: Kuai et al., arXiv 2604.07650. Table C.2 (top 10 BEI with q-values), Table C.3 (CIG), §B.1 roster and Instruct naming convention, results prose and Appendix C text on intra-family entanglement.
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels: Kohli, arXiv 2605.29800. φ̄ = 0.391, n_eff = 2.18, 22.0pp Condorcet gap under human-entropy conditioning.
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus: Hossain, Yousefi & Lim, arXiv 2609.22512. Cross-provider 0.42 vs within-Gemini 0.40; difficulty adjustment 0.465 → 0.467.
- Item Response Theory for LLM Benchmarks and Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks: ATLAS, arXiv 2511.04689. Table 4 split-half and cross-benchmark rows, the low-ability compression band, the temporal holdout.
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?: Zheng & Yang, arXiv 2609.00482. Family DIF replication, K = 8–128, near-tie reversal and matched-random rates, claim boundaries (no future-item estimation; no release-date control).
- Economic Benchmark Construct Validity and One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation: Zhu, arXiv 2608.29420. Release-date R² 0.505 and the leave-one-benchmark-out predictor ladder.
- Benchmark Score Redundancy: Zhu's confound on the score matrix, and CollabEval's item-level dimensionality.
- Benchmark Convergent and Discriminant Validity: Desai et al., 1PL ΔAUC as a dimensionality diagnostic.
- Version-Dependent Judge Error and Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation: Li, arXiv 2609.34198. Capability gradient in judge false acceptance (+0.82).
- Cognitive Capability Profiling for Task Suitability: Prunty et al. Rubric-sourced difficulty; no external outcome validation.
- Recurring Production Agent Evaluation: She & Lin. MD-2PL vs Rasch; IRT reconstruction as within-benchmark prediction.
Cited by 4
- Cross-Model Error Entanglement×3
Construct Or Roster Latent Structure Readings: reads BEI's top tier as a six-model tier clique in…
- Item Response Theory for LLM Benchmarks×3
Construct Or Roster Latent Structure Readings: sets the θ̂-validity question beside BEI's vintage…
- Open Questions Backlog×2
Cross Model Error Entanglement: Is the BEI graph measuring entanglement or vintage? → Construct Or…
- Evals & Benchmarks
Construct Or Roster Latent Structure Readings — Two #oq/now items answered as one: both read shared…
Related articles
- Benchmark Score Redundancy
Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix comp…
- Item Response Theory for LLM Benchmarks
Score a benchmark with a 3PL IRT model instead of percent-correct and two separable things follow. Scoring: ability θ r…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Evals & Benchmarks
Map of Content for the evals-and-benchmarks domain — 36 concepts. The science of measuring models: benchmark validity,…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
