Sources#
- A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Summary#
Every multi-model construction on this wiki — a judge panel, a verifier ensemble, a maker/checker split, a majority vote — buys its reliability from an assumption of independent errors, and LLM-Judge Validation, Reference-Free Judge Over-Crediting and Weak-Verifier Ensembling each carry a different measurement saying that assumption is false. What none of them carries is an instrument whose estimand is the dependence itself, measured between the models rather than between their verdicts, with a null model and a significance test attached.
Kuai et al. (A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges, arXiv 2604.07650, v2 2026-08-09, Texas A&M / Marquette / Utah, COLM 2026, empirical) supply one. The framing sentence is the paper's own:
correct answers naturally converge because tasks typically have a single solution. In contrast, errors occupy a much larger hypothesis space.
So the audit is run only on the failure manifold. Two models agreeing on a right answer is uninformative; two models failing the same easy question, and failing it toward the same wrong option, is not.
This page now holds two such instruments, pointed at different objects and arriving at the same verdict. Kuai et al. measure dependence between models as answerers and predict a judge's bias from it. Kohli (Apple, arXiv 2605.29800) measures it between a judge panel's own verdicts and converts it into the accuracy the panel is losing: nine frontier judges from seven families carry n_eff = 2.18 independent votes, and majority voting lands 22.0pp below what independent voting with the same per-judge error profiles would achieve. Read the two together and the shape is: the dependence is real on every instrument, it is not explained by item difficulty on either, and no aggregation rule yet tested recovers more than a tenth of what it costs.
What is being measured, and on what#
The title says both "behavioral dependence" and "LLM judges", and the two are measured on different objects. This distinction decides what the page's numbers can be used for:
- Dependence is measured on answering behaviour, not on judging behaviour. All 18 models answer a 1,000-question MMLU-Pro subset (biology/psychology, 473/527) under a unified 5-shot prompt at temperature 0 where supported. BEI and CIG are computed pairwise from that error matrix. No judge verdict enters.
- Bias is measured on judging behaviour, on a disjoint set. Three of the 18 — Llama-3.1-70B-Instruct, GPT-5, GPT-4o-mini (open-source, high-capability closed, cost-efficient closed) — then judge candidate responses on a second, disjoint 1,000-question MMLU-Pro subset (law/business, 295/705), scoring correctness (0/1) and reasoning quality (0–5) against a fixed rubric, with all candidate responses in one shuffled prompt.
- The join is a rank correlation across (judge, target) pairs, not a mechanism. The dependence estimated in the first stage predicts a bias quantity measured in the second. Nothing propagates from one to the other by construction, which is what makes the transfer a result rather than an identity — and also what makes it an association only.
The paper is explicit that "induced" in its title is not a causal claim: "the proposed audit identifies behavioral association but does not establish its causal origin." Read every number below as correlational.
Level 1 — BEI, excess co-failure conditioned on difficulty#
The naive version of this measurement fails for an obvious reason: two weak models co-fail constantly because hard questions are hard. BEI conditions that away.
For models i, j on task t, the difficulty is estimated leave-pair-out — the failure rate among the other M − 2 models, d_t = (1/(M−2)) Σ_{m ∉ {i,j}} Y_{m,t} — so the pair under test never enters its own conditioning variable. Each model's difficulty-response function p_m(d) = P(Y_m = 1 | d) is fitted by logistic regression, and the null is conditional independence given difficulty, H₀: Y_{i,t} ⊥ Y_{j,t} | d_t, under which the expected co-failure is p_i(d_t)·p_j(d_t).
The residual Y_{i,t}Y_{j,t} − p_{i,t}p_{j,t} is then weighted by task easiness a_t^w = (1 − d_t)^w:
BEI_w(i,j) = (1/T) Σ_t (1 − d_t)^w (Y_{i,t}Y_{j,t} − p_{i,t}p_{j,t})
The weighting is the design decision and the paper argues it well: on a hard task the null co-failure probability is already high, so a non-co-failure produces a large negative residual that means nothing much (one model was simply capable); a co-failure on an easy task, where both models were individually expected to succeed, is the diagnostic event. Proposition 1 shows the easiness weight is a deterministic function of d_t and so preserves the null centring at zero — the result holds task-wise and does not require independence across tasks.
Significance is Monte Carlo: resample both models' outcomes as independent Bernoulli draws at the fitted p̂_{i,t}, p̂_{j,t} with observed difficulties held fixed, recompute BEI, take a one-sided p-value against the simulated null, then Benjamini–Hochberg across all model pairs and report q-values. The difficulty-response fits are diagnosed rather than assumed: mean AUC 0.865 (range 0.752–0.934), mean ECE 0.041, mean Brier 0.130 (Table C.1).
Level 2 — CIG, excess same-distractor collision#
BEI says two models fail together more than chance. CIG asks whether they fail the same way. On an MCQ task with K − 1 distractors, distractor attractiveness π_{k,t} is estimated with add-one smoothing from the failing responses of the other models (pair-excluded again), and the directional null is that two co-failing models draw their distractors independently from that distribution, so the null collision probability is c_t = Σ_k π²_{k,t}. Each collision residual is weighted by its own surprisal −log c_t:
CIG(i,j) = (1/T) Σ_{t ∈ co-failures} (−log c_t)(1[S_i = S_j] − c_t)
A collision on a task where one distractor is a near-universal trap counts for little; a collision on a task where the other 16 models scattered counts for a lot. Same Monte Carlo + BH machinery.
The pair divides cleanly, in the paper's words: Level 1 captures whether two models excessively fail together, Level 2 whether they excessively fail in the same way.
What the audit found#
The roster (Appendix B.1 prose, cross-checked against the 18 nodes of Figure 2) is 18 models, six families, and it deliberately mixes vintages:
| Family | Models |
|---|---|
| GPT | GPT-5, GPT-4o, GPT-4o-mini, GPT-oss-20B |
| Claude | Claude 4.6 Sonnet, Claude-3.5-Sonnet, Claude-3.5-Haiku |
| Gemini | Gemini-3.1-Pro, Gemini-1.5-Pro, Gemini-1.5-Flash |
| Llama | Llama-3.1-70B, Llama-3.1-70B-Instruct, Llama-3-70B, Llama-2-70b |
| Qwen | Qwen1.5-110B, Qwen1.5-72B-Chat, Qwen1.5-14B-Chat |
| DeepSeek | DeepSeek-Chat-v2.5 |
The two metrics find structurally different graphs, and that is the most useful thing in the paper.
- BEI's top 10 pairs are entirely Llama and Qwen (Table C.2, reconciled against
pdftotext -layout): Llama-3-70B ↔ Llama-3.1-70B at 0.0525, Llama-2-70b ↔ Qwen1.5-110B at 0.0398, Llama-3-70B ↔ Qwen1.5-110B at 0.0352, Qwen1.5-14B ↔ Qwen1.5-72B at 0.0333, down to 0.0211, every one significant after FDR (q = 9.5 × 10⁻⁴ to 1.3 × 10⁻²). Not one frontier model appears. Co-failure entanglement, as measured, is a property of the older open-weight tier. - CIG's top pairs are where the cross-family story lives (Table C.3, nine rows, reconciled): Llama-3-70B ↔ Llama-3.1-70B still leads at 0.1332, but then Claude-3.5-Sonnet ↔ GPT-4o (0.0525), GPT-4o ↔ GPT-4o-mini (0.0523), DeepSeek-v2.5 ↔ Gemini-1.5-flash (0.0502), Claude-4.6-Sonnet ↔ GPT-5 (0.0471), Gemini-1.5-flash ↔ Gemini-1.5-pro (0.0454), Claude-3.5-haiku ↔ Gemini-1.5-flash (0.0429). Directional error alignment crosses vendor boundaries that co-failure does not.
- The prose adds a compositional observation the tables do not carry: "more statistically significant entanglement pairs observed among closed-weight models and fewer between open- and closed-weight models." That is the closest thing in the corpus to a directional answer on what provider-mixing buys — mixing across the open/closed line decorrelates more than mixing within a tier — and it is stated as a count of significant pairs, with no effect size.
The authors' own reading: entanglement "extends beyond architectural lineage or model genealogy" and may reflect "shared training signals, alignment strategies, and generation-era design choices" — with the causal disclaimer attached every time they say it.
The judge-bias association, and what its arithmetic will bear#
Over-endorsement is operationalised as a deviation in conditional precision:
ΔPrec(J_j, M_i) = P(Y=1 | Ŷ_j=1) − P(Y=1 | Ŷ_j=1, M=M_i)
— how much lower the judge's precision is for this particular target model than its precision overall. Positive means the judge endorses that model's wrong answers at an above-average rate. Spearman correlations against the dependence metrics (Table 1, reconciled cell-for-cell against pdftotext -layout):
| Metric | MMLU-Pro ρ | p | MATH-500 ρ | p |
|---|---|---|---|---|
| Pearson corr. | 0.063 | 0.724 | 0.187 | 0.360 |
| Mutual information | 0.113 | 0.525 | 0.208 | 0.309 |
| Co-failure Jaccard | 0.428 | 0.012 | 0.275 | 0.175 |
| BEI | 0.508 | 0.002 | 0.441 | 0.024 |
| Error-only agreement | −0.107 | 0.545 | −0.172 | 0.399 |
| Difficulty-adj. residual corr. | 0.151 | 0.394 | −0.057 | 0.746 |
| CIG | 0.520 | 0.002 | 0.457 | 0.019 |
The MATH-500 column is the paper's strongest structural move: dependence estimated on multiple-choice MMLU-Pro, applied unchanged to free-form mathematical responses judged by the same three judges on the 14 models still available at evaluation time, and it still predicts. CIG is not even re-estimated there — it is defined for MCQ and the directional structure is transferred, not recomputed. Most baselines go non-significant on that transfer; BEI and CIG do not.
Three cautions, and the first is load-bearing.
1. The correlations run on ~34 and ~26 points, and the paper never says so. No sample size for any Spearman appears anywhere in the text or the tables. Back-solving each column from its own (ρ, p) pairs under the usual t-approximation gives n = 33–34 for every MMLU-Pro row and n = 26 for every MATH-500 row — recovered consistently from all seven rows in each column, including the null ones (ρ = 0.063 / p = 0.724 and ρ = 0.187 / p = 0.360 both land on the same n as the significant rows), so it is a single shared n and not an artefact of one cell. Wiki arithmetic, not the paper's. Two consequences:
- The confidence intervals are wide enough to swallow the paper's central contrast. At n = 34, BEI's ρ = 0.508 carries a 95% interval of roughly [0.21, 0.72] and CIG's 0.520 [0.22, 0.73], while mutual information's 0.113 carries [−0.23, 0.44] and Jaccard's 0.428 [0.11, 0.67]. "BEI and CIG show the strongest associations… in contrast, Pearson correlation and mutual information exhibit only weak and non-significant relationships" is a comparison between overlapping intervals. The sign and significance of BEI and CIG survive; their superiority over the baselines is not established at this n. The honest version of the claim is the MATH-500 column, where Jaccard drops to non-significance and BEI/CIG do not — a difference in transfer, not in magnitude.
- The accounting does not close. Three judges × 17 remaining targets is 51, which would have produced p ≈ 1.4 × 10⁻⁴ at ρ = 0.508, not the reported 0.002. The numbers that fit exactly are 2 × 17 = 34 and 2 × 13 = 26 — two judges, not the three the setup describes. Either one judge is excluded from the headline association or the unit of observation is something the paper does not state. It should have been one sentence in the Metrics paragraph.
2. The difficulty-weighting exponent w is a free parameter chosen on the reported data. Table 2 (reconciled): at w = 0, where BEI degenerates to the plain mean co-failure residual, ρ falls to 0.289 (MMLU-Pro) and 0.032 (MATH-500) — so essentially all of the metric's advantage over a raw residual comes from the easiness weighting, which is the paper's own point and is well made. But the transfer degrades fast above w = 1: MATH-500 ρ runs 0.441 → 0.275 → 0.231 → 0.236 → 0.180 at w = 1…5 while MMLU-Pro holds at 0.508 / 0.490 / 0.504 / 0.465 / 0.363. The conclusion that "moderate weighting, around w = 1–2, provides a reasonable balance" is drawn from the same table that reports the headline; on the cross-benchmark axis only w = 1 actually works, and no held-out selection of w is performed.
3. A trivially cheap baseline already gets most of the way on the primary benchmark. Co-failure Jaccard — set overlap of failures, no difficulty conditioning, no null model, no inference — reaches ρ = 0.428 (p = 0.012) on MMLU-Pro. The machinery earns its keep on transfer and on the directional half, not on the primary correlation.
The mitigation, and the delta that is actually attributable to it#
The downstream demonstration reweights a three-verifier ensemble. Each verifier J_m gets a competence score q_m and two dependence penalties — R_m, its mean entanglement with the other verifiers, and T_m^(S), its entanglement with the target model being judged — combined as
w_m^(S) ∝ q_m^κ · exp(−η₁ R_m − η₂ T_m^(S))
with λ mixing normalised BEI and CIG into the pairwise entanglement E(i,j). All four parameters are fitted by cross-entropy on a 500-question calibration split and frozen before the 500-question held-out split. That discipline is real and worth crediting.
Results (Table 3, reconciled):
| Aggregation | Acc | F1 | Precision |
|---|---|---|---|
| Majority vote | 0.847 | 0.901 | 0.870 |
| Accuracy-based reweight | 0.881 | 0.919 | 0.891 |
| Entanglement-based reweight | 0.882 | 0.922 | 0.896 |
The abstract's "3.5 and 2.6 percentage-point gains in accuracy and precision… over majority voting" is arithmetically correct and rhetorically misplaced. The competence-only baseline already delivers +3.4pp of the +3.5pp. The increment attributable to the entanglement term — the paper's actual contribution to this table — is +0.1pp accuracy, +0.3pp F1, +0.5pp precision, on a 500-question held-out split where one point of accuracy is roughly one standard error. Figure 3 shows why: the calibrated weights are dominated by competence (GPT-5 sits near 0.42–0.50 across every target model, GPT-4o-mini near 0.25, Llama-3.1-70B-Instruct in between) and move only slightly as the target changes. The de-entangling adjustment is visible in the heatmap and invisible in the metric.
This is the honest summary of the mitigation half: the audit is the contribution; the reweighting is an existence proof that the audit's output can be plugged into an aggregator without hurting, on a three-verifier pool too small for a redundancy penalty to have much to work with. It does not yet show that pricing dependence buys anything over pricing competence.
Nine judges, two effective votes — the same deficit priced on the verdicts (2026-09-22)#
Kohli (Apple, Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels, arXiv 2605.29800, 2026-05-28, 14pp, empirical) runs the measurement the sections above cannot: dependence estimated on a judge panel's own verdicts, while the panel does its actual job, against a ground truth rich enough to price the consequence in accuracy points rather than in a correlation coefficient.
The setup, and why the dataset choice is the whole design#
Nine frontier judges from seven families — GPT-4o and GPT-4o-mini (OpenAI), Claude Sonnet 4.5, Gemini 2.5 Pro, Llama 4 Maverick and Llama 4 Scout (Meta), Qwen3-32B, Mistral Large 3, DeepSeek-V3 — classify NLI items at temperature 0 under one standardised prompt. Individual MNLI error rates run 0.282 (Qwen3-32B, best) to 0.356 (GPT-4o-mini) (Table 1, reconciled).
The corpus is ChaosNLI (Nie et al., 2020): 100 human annotations per item. Gold is the 100-annotator majority; 1,000 entropy-stratified items are drawn from each of MNLI (1,599 available), SNLI (1,514) and AlphaNLI (1,532), plus 1,000 RewardBench pairwise-preference items as a fourth task. That annotation depth does two jobs at once — it supplies an answer key and a per-item difficulty measure (human Shannon entropy, 0–1.58 bits on the 3-class sets, 0–1.00 on AlphaNLI) — and the second job is what lets the paper condition dependence on difficulty without borrowing difficulty from the models under test, which is the leave-pair-out contortion Kuai et al. need.
The instrument: Kish effective sample size#
Each judge gets a binary error vector over the 1,000 items (e_{j,i} = 1 if judge j disagrees with the ChaosNLI majority label); the phi coefficient is computed across all C(9,2) = 36 pairs; the Kish design effect converts the mean into an effective count of independent voters:
n_eff = k / (1 + (k − 1) · φ̄)
Phi is chosen deliberately: for binary error vectors it is the Pearson product-moment correlation the Kish formula requires, whereas "alternative association measures (e.g., Cohen's kappa) conflate prevalence with dependence."
Headline (Table 2, reconciled against pdftotext -layout): φ̄ = 0.391 ± 0.111 (pair range 0.161–0.603), n_eff = 2.18 [2.07, 2.31] (bootstrap over 10,000 item resamples), independence ratio n_eff / k = 24.2%. The exchangeability-free check agrees — k/λ_max = 2.16 at λ_max = 4.17 — so the Kish assumption holds for this panel. Nine judges, two votes.
The consequence: a 22-point Condorcet gap that difficulty does not explain#
n_eff on its own is a statistic about a correlation matrix. The paper converts it into a claim about the panel's output with a Condorcet null: per-judge 3×3 confusion matrices estimated per human-entropy tercile, then 10,000 Monte Carlo draws per item in which each judge votes independently from its own bin-specific confusion matrix given the item's gold label. This is explicitly a conditional independence test — correlation that shared item difficulty explains is given away for free, and the gap is what survives it.
- Predicted majority-vote accuracy under independence: 94.0%. Actual: 72.0%. Weighted gap 22.0pp [19.5, 24.1].
- Only 6.8% of the gap is attributable to shared item difficulty (13.5% with 10 bins instead of 3; 66–87% remains unexplained across all datasets).
- Stratified permutation test (10,000 permutations shuffling each judge's error vector within entropy stratum, preserving per-judge error rates and the difficulty structure): 0 of 10,000 permutations reached the observed φ̄. Null mean 0.060, SD 0.005, z = 65.6,
p < 10⁻⁴. - Split-half cross-validation (confusion matrices fitted on one 500-item half, simulated on the other) returns 21.9pp against the in-sample 22.0pp — overfitting ratio 0.997, and 0.960/1.000 on SNLI/AlphaNLI.
The error histogram is the picture of it (Figure 1, read at compile time; its ten bars sum to 1,000 — 290/140/103/108/68/69/65/61/45/51). Under independence the mass concentrates at 2–4 errors per item. Observed, it piles up at both ends: 290 items (29%) with all nine judges correct, 51 (5.1%) with all nine wrong against fewer than 1 expected for the all-wrong cell.
Three results that a correlation coefficient alone would not have produced#
1. The best single judge matches or beats the panel — on every dataset. Table 3 (cross-dataset, the one table in this raw that needed repair — see Sources):
| MNLI | SNLI | AlphaNLI | |
|---|---|---|---|
| n_eff (Kish) | 2.18 [2.07, 2.31] | 2.35 [2.21, 2.51] | 2.48 [2.32, 2.69] |
| φ̄ | 0.391 | 0.354 | 0.328 |
| Panel accuracy | 72.0% | 77.7% | 88.7% |
| Best individual | 71.8% | 84.2% | 91.2% |
| Panel lift | +0.2pp | −6.5pp | −2.5pp |
| Condorcet gap | 22.0 [19.5, 24.1] | 14.0 [11.9, 16.1] | 7.6 [6.0, 9.1] |
The +0.2pp on MNLI is inside the tie-breaking margin (11 hash-broken ties, 1.1% of items), so the honest reading is that the panel never wins. Note the gap shrinks with base accuracy — higher accuracy leaves less room for correlated error — which is why the RewardBench gap is only 6.8pp at 92.7% panel accuracy while φ̄ there is higher, 0.440. A small Condorcet gap is not evidence of an independent panel.
2. Adding judges is nearly free of value, and removing them is often positive. The scaling curve (Table 8, C(9,k) subsets) tracks the Kish prediction almost exactly — mean n_eff 1.45 / 1.69 / 1.85 / 1.96 / 2.03 / 2.09 / 2.14 / 2.18 at k = 2…9 against predicted 1.44 / 1.68 / 1.84 / 1.95 / 2.03 / 2.09 / 2.14 / 2.18 — with a hard asymptote at 1/φ̄ = 2.56. The first five judges buy 90% of the achievable independence (1.96 of 2.18); judges 6–9 buy +0.22 effective votes. Leave-one-out (Table 9) is worse than flat: six of nine removals raise panel accuracy, and removing Gemini 2.5 Pro — the judge most entangled with Claude (φ = 0.603) and GPT-4o (0.52) — raises it by 1.3pp [+0.1, +2.6]. The three removals that hurt include the two most individually accurate judges. Austen-Smith and Banks (1996) predicted that voters can hurt under positive correlation; this is the demonstration in the LLM-judge setting.
3. Smarter aggregation does not rescue it, including with oracle labels. Table 5 / Table 11 (reconciled), accuracy in %:
| Method | Oracle? | MNLI | SNLI | AlphaNLI | RewardBench |
|---|---|---|---|---|---|
| Majority vote | no | 72.0 | 77.7 | 88.7 | 92.7 |
| Dawid–Skene EM | no | 70.7 | 77.6 | 89.5 | 92.7 |
| Accuracy-weighted (5-fold CV) | yes | 72.2 | 77.7 | 88.7 | 92.7 |
| Phi-optimal / Markowitz (5-fold CV) | yes | 72.4 | 78.4 | 86.2 | 94.1 |
| Best individual judge | — | 71.8 | 84.2 | 91.2 | 95.5 |
| Condorcet prediction | — | 94.0 | 91.7 | 96.3 | 99.5 |
Accuracy-weighted voting closes less than 1% of the MNLI gap. Dawid–Skene underperforms majority vote on MNLI (unsupervised EM misestimates error rates when judges are correlated) and its best result anywhere is 10.5% of the AlphaNLI gap. The stable oracle-informed methods close at most 11%. Markowitz weighting — the direct "price the correlation" move, inverting the phi matrix exactly as Weaver-style weighting would like to — is the most instructive failure: best on MNLI (72.4) and RewardBench (94.1, closing 20.6% of that gap) and below plain majority vote on AlphaNLI (86.2). Fitting weights to a correlation structure overfits it.
Same-family versus cross-family: the lever the corpus keeps assuming#
This is the result that bears hardest on the rest of the wiki, and the phi matrix (Figure 3, read at compile time and cross-checked against the §5.3 and Appendix B prose) settles it visually: the hierarchically clustered matrix has a tight five-model block — Mistral Large, DeepSeek-V3, GPT-4o, Claude Sonnet, Gemini 2.5 Pro, pairwise 0.46–0.60 — that spans five different vendors, and the loose corner is the two Metas plus Qwen3 and GPT-4o-mini. The OpenAI pair is split across the two clusters (GPT-4o inside the tight block at 0.59 with Claude and 0.52 with Gemini; GPT-4o-mini outside it), which is the clustering's own statement that vendor is not the partition.
- Same-family pairs: OpenAI×OpenAI φ = 0.437, Meta×Meta 0.435. Cross-family mean φ̄_cross = 0.389. Difference: +0.047.
- The three most correlated pairs in the panel are all cross-family: Claude Sonnet × Gemini 2.5 Pro 0.603, GPT-4o × Claude Sonnet 0.588, Mistral Large × DeepSeek-V3 0.564. The least correlated pair is also cross-family (Gemini × Llama 4 Scout, 0.161).
- Maximising family diversity makes it worse, not better. Restricting the panel to one judge per family — seven judges, the best in each — gives n_eff = 1.93, below the 9-judge panel's 2.18 and below the mean of all 36 random seven-judge subsets, which Table 8 puts at 2.09 (wiki arithmetic on the paper's own numbers). The paper's reading is a selection effect: the best judges concentrate their errors on the same hard items. Either way, "one per vendor" is a panel-construction rule that buys negative independence here.
- One counterweight worth carrying: the same-family excess is +0.109 on RewardBench against +0.047 on MNLI. Family is a stronger proxy on pairwise preference than on classification, so "lineage doesn't matter" is task-dependent rather than universal.
Unanimity is not what it looks like#
Among the 319 unanimous MNLI items the panel is 90.9% accurate; on the 681 items with any disagreement, 63.1%. A 9.1% error rate on unanimity sounds unremarkable until it is set against what independence predicts for nine voters at ~68% accuracy: ~0.02%. The per-bin table (Table 6, 11 rows) is significant at p < 0.05 in 8 of 12 discrete panel-entropy levels. And the all-wrong items are not merely the ambiguous ones — of the 51, 29 (56.9%) are items where at least 50% of humans agreed on the answer the panel missed, and over-prediction of contradiction accounts for 51% of them (n→c 14, e→c 12) despite contradiction being only 16.5% of gold labels.
This is the quantitative form of a rule several pages here already state informally: panel agreement is a measurement of shared prior, not of correctness, and using unanimity as a confidence gate imports the correlation wholesale.
Robustness, and the one knob that makes it worse#
Re-running all nine judges on the same 1,000 MNLI items (Table 4): reframed prompt 2.17, reversed label order 2.15, temperature 0.5 2.17, against baseline 2.18 — overlapping CIs throughout, and the reversed-label arm rules out position bias as the driver. RewardBench, a different task with deterministic gold labels, lands at 1.99 [1.83, 2.20]. Stratifying MNLI by gold class, n_eff runs 1.85 (contradiction) to 2.40 (neutral); even on the 179 items where ≥80% of humans agree it is 2.67 — higher than the full set, and still nowhere near 9, which kills the "herding is just ambiguity" reading.
The exception is the knob practitioners reach for first. Chain-of-thought raises correlation: φ̄ 0.391 → 0.456, n_eff 2.18 → 1.94, panel accuracy 72.0% → 69.2%. Shared reasoning amplifies shared error. Anyone buying panel diversity with "make the judges think" is buying the opposite.
The recommendation, and what it costs to follow#
We recommend computing n_eff as a standard panel diagnostic: if n_eff/k < 0.5, results should be treated with caution.
Every configuration in this paper fails that test (24.2% at k = 9; 1.99/9 = 22% on RewardBench). The diagnostic is cheap where an answer key exists — a panel's own error matrix is a by-product of any scored eval run — and the progress metric is explicit: halving φ̄ from 0.39 to 0.20 raises n_eff from 2.2 to 3.5 and closes roughly half the Condorcet gap.
Three cautions travel with it. The human comparator is simulated, not measured. Table 3's "human n_eff" of 4.03–5.79 — the paper's basis for "human panels are about 2× more independent" — is obtained by sampling 10 labels per item from the aggregate ChaosNLI distribution under an exchangeability assumption (the paper's own footnote), so it is what an exchangeable annotator pool would score, not a measured n_eff for the 100 real annotators. The bootstrap CI covers items, not judges — the paper says so; a different nine-model panel could land elsewhere. And everything here is classification or binary preference; the authors name open-ended generation and code review as the untested regimes, which is exactly where judge panels are actually deployed.
Where the two papers agree, and the one place they look like they don't#
They agree on the thing that matters: errors are not independent, the deficit survives conditioning on difficulty, and a null model is what makes that statement checkable. They also agree on the mitigation's ceiling from opposite directions — Kuai et al. build the de-entangled aggregator and recover +0.1pp over competence weighting; Kohli tests the two nearest established aggregators and recovers ≤11% of the gap with oracle labels.
The apparent conflict is on family. Kuai et al.'s BEI top-10 co-failure pairs are entirely intra-family (Llama and Qwen); Kohli's top-3 verdict-correlation pairs are entirely cross-family. The constructs resolve it:
- BEI measures excess co-failure occurrence, difficulty-conditioned, on answering. That is where lineage shows — and on Kuai's roster it is confounded with vintage, since the whole top tier is 2023–2024 open-weight models.
- CIG measures directional error alignment, and its top pairs cross vendors (Claude-3.5 × GPT-4o, DeepSeek × Gemini-flash). Kohli's φ is a verdict-level correlation, which is closer to CIG's construct than BEI's, and it agrees: alignment crosses lineage.
- Kohli's panel is the single-release-window control that Kuai's roster lacks. All nine judges are current frontier endpoints, no three-generation spread, and the cross-family result holds anyway at φ̄ = 0.391. That is an independent line of evidence against the vintage reading of BEI's intra-family tier — on a different instrument and task, so not a replication, but pointing the same way.
One construct difference to keep straight: Kohli's φ is unconditioned — the raw pairwise correlation of error vectors, which is the plain baseline Kuai et al. find predicts judge bias only weakly. Kohli does not use it to predict bias; he uses it inside Kish, where the raw correlation is the right input, and he applies the difficulty conditioning one level down, in the Condorcet null. The two papers put the same correction in different places, and both find the same residual.
A third n_eff measurement, and a new consequence: pooled votes vs. paired significance (2026-09-25)#
Hossain, Yousefi & Lim (Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus, UCF, arXiv 2609.22512, 2026-09-18, empirical) run the same design-effect measurement as Kohli — Ledoit-Wolf-shrunk error correlation → n_eff = n / (1 + (n−1)ρ̄) — on a different bank, a different ground-truth source, and a different downstream consequence. It is the third independent group in the corpus (after Kuai et al. and Kohli) to build a null model for judge/model error dependence rather than report a raw agreement number, and the three groups' methods do not overlap: Kuai conditions on task difficulty and needs no judge ground truth beyond the answer key, Kohli anchors n_eff to a Condorcet accuracy prediction, and this paper anchors it to a hypothesis-test conclusion.
The headline triangulates Kohli's, on a differently-composed bank. Ten open-weight logical judges (five 7–9B base models × pairwise/Likert-then-compare prompting: Llama-3.1-8B, Qwen-2.5-7B, Gemma-2-9B, Mistral-7B, Phi-3.5-mini) on 1,133 listwise-complete RewardBench v2 pairs give ρ̄ = 0.206, n_eff = 3.51 of 10 (Table 9) — a 35% independence ratio, close to Kohli's 24.2% on a smaller, frontier-heavier panel, and the same qualitative story: nominal panel size overstates independent evidence by roughly 3×. The design-effect variance inflation is 1 + 9·0.206 ≈ 2.85. The estimate is stable to calibration-label noise (ρ̄ moves only 0.206 → 0.226 as 1–10% of gold labels are randomly flipped) and generalizes across 73 alternative bank compositions (ρ̄ ranging 0.13–0.28 across 68 subbanks of ≥5 judges) and two more datasets (n_eff = 3.36 on UltraFeedback, 2.89 on PKU-SafeRLHF) — the broadest robustness sweep behind any single n_eff estimate in the corpus.
The new consequence: correlated votes flip which significance test says "significant." Where Kohli converts n_eff into a predicted accuracy gap, this paper converts it into a predicted hypothesis-test disagreement. A "pooled-vote" test that treats every judge vote on every item as an independent observation understates the variance of the mean vote by the same design-effect factor (≈2.85); a "pair-based" test that treats each preference pair as the unit of inference does not. On 5,000 resampled 100-pair panels drawn from RewardBench Focus and Factuality subsets, the pooled test is significant in 79% of panels against the pair-based test's 52% — so in up to 28% of panels (22% on Factuality alone, ≤2% on Math and Precise IF) the two tests disagree about whether one system beats another, using the same 1,000 votes: one worked example gives z = 2.03 pooled vs. z = 1.52 paired. This is the sharpest form yet in the corpus of the item-as-unit-of-inference argument Yang et al. and Kohli both make from the accuracy side — here it is a conclusion that reverses, not a number that shrinks, and it needs no ground truth beyond the preference pairs themselves (RewardBench's own reference labels), which makes it cheaper to run than either accuracy-based instrument.
Frontier judges triangulate Kohli's cross-family finding on an entirely fresh roster. Three different-provider frontier judges — GPT-5.6-sol, Claude Opus 5, Grok 4.5 — at 91–93% RewardBench accuracy have ρ̄ = 0.56 [0.42, 0.68], n_eff = 1.41 of 3, and co-fail together at 7.7× the rate independence predicts (Pr[B fails | A fails] = 0.64 against a 0.083 marginal). Decomposed by pair class (Table 8), cross-provider frontier pairs correlate at 0.42, essentially the same as the 0.40 measured between three Gemini models from one provider — a second, independent instance of Kohli's finding that vendor is not the decorrelating variable (Kohli: same-family 0.437/0.435 vs. cross-family mean 0.389, a +0.047 gap in the other direction but on the same order of magnitude; here the gap runs the opposite sign, cross ≥ within, on a different task and a newer judge generation). Adjusting for item difficulty measured by the open-weight bank leaves the frontier-pair correlation almost unchanged (0.465 → 0.467 on RewardBench), so shared difficulty does not explain it either — the same "not vintage, not difficulty" move Kohli and the partially-answered open question below already established, now repeated on GPT-5.6-sol/Opus 5/Grok 4.5 rather than GPT-4o/Sonnet 4.5/Gemini 2.5 Pro.
A mechanism test neither Kuai nor Kohli ran: does preference optimization cause the frontier-level dependence? Training one GRPO QLoRA adapter per base model (three open-weight models, matched forced-choice protocol, identical items and candidate positions before and after) gives Δρ̄ = −0.021 [−0.035, −0.007] — GRPO fine-tuning slightly lowers error correlation, not raises it, while mean accuracy moves only 0.7pp. DPO shows no significant effect ([−0.014, +0.002]). Both trained banks stay far below the frontier banks' 0.32–0.60 range. The result is explicitly scoped (100 GRPO steps, one seed, three base models, not fully held out from the calibration data used for the adapters), but it is the first attempt in the corpus to test a specific causal hypothesis for why frontier judges are more correlated, and it comes back negative for the simplest candidate mechanism: preference optimization alone does not reproduce or explain frontier-level dependence.
A regime distinction this paper adds that neither Kuai nor Kohli make. ρ̄ and n_eff can move in opposite directions depending on how errors are distributed: when a vulnerable subgroup of judges fails together on position-sensitive items, injecting more of that failure lowers measured bank-wide correlation (ρ̄ 0.22 → 0.08, apparent n_eff rising 3.4 → 6.0) even as majority-vote precision falls — "the bank appears more independent even as its majority decisions become less reliable." The regime-matched filtering this enables (CorrFilter for shared-error "global co-failure," Bias-Cluster for concentrated "vulnerable-subgroup failure") is developed in full on Weak-Verifier Ensembling, since it is a dependence-aware aggregation method rather than a dependence measurement.
Three instruments, three different dependences#
The corpus now holds four ways of asking whether models' errors are independent, and they are not interchangeable. Getting the constructs apart is most of the value of having all four.
| Measured on | Needs ground truth | Needs an external anchor | Identified? | |
|---|---|---|---|---|
| Yang's intra-class ρ (LLM-Judge Validation) | judge verdicts, from a vote matrix | no | no | yes, but it cannot say what the jurors agree about |
| Sunkavalli's ρ_k (Weak-Verifier Ensembling) | judge scores, as covariance | no | yes, ≥ 2 anchors | only under continuous scores and no panel-wide residual |
| BEI / CIG (this page) | model answers, on the failure manifold | yes | no | yes, by construction |
| Kohli's Kish n_eff (this page) | judge verdicts, from the panel's own error matrix | yes (100-annotator majority) | no | yes for the one job it has — turning φ̄ into a predicted majority-vote accuracy |
The trade is exact and worth stating plainly. Sunkavalli's identification wall exists because judge covariance K = σ_t² + σ_c² fuses signal with shared error, and no within-panel statistic splits them — high agreement is equally consistent with a panel that is right and a panel that is uniformly wrong. BEI and CIG do not hit that wall, because ground truth splits it for them: conditioning on Y = 1 discards every agreement-because-correct, so what remains is shared error by construction, and the difficulty-response fit supplies the null that the covariance approach has to estimate. The price is that the method only works where objectively verifiable answers exist — the paper says so in its first experimental sentence — which excludes the open-ended generation tasks where judge panels are most used and where the shared-latent-plausibility failure of Reference-Free Judge Over-Crediting actually bites.
The second difference is the object. Yang and Sunkavalli measure dependence between judges. Kuai et al. measure dependence between models as answerers, then use it to predict how those models are judged. That is why this page can speak to Same-Model Review Blindness and Agent Behavioral Homogeneity at all, and why its verifier-reweighting numbers are weaker than its audit numbers: the reweighting needs the answerer-side dependence to be a good proxy for verifier-side dependence, and nothing here tests that step.
The fourth row is the one the corpus was missing, and it is now filled. Yang's ρ prices a jury from its own votes but cannot say what the jurors agree about; Sunkavalli's ρ_k can separate signal from shared error but needs external anchors and dies under ordinal scores; BEI/CIG buy identification with ground truth but only where verifiable answers exist, and measure models as answerers. Kohli's n_eff measures the panel in situ — real judges, real verdicts, no anchors — and pays for it by needing a ground truth as expensive as ChaosNLI's 100 annotations per item, and by measuring the raw correlation with no null of its own (the null lives downstream, in the Condorcet simulation). It is the only one of the four that converts dependence directly into the number a practitioner cares about: how many accuracy points the panel is losing.
Limits the paper states, and one it does not#
Stated: the framework is pairwise only (no higher-order dependence); CIG is defined for MCQ, so MATH-500 tests transfer rather than re-estimation; the audit identifies association, not causal origin; experiments are English-only; and — the line that should travel with every number here — "because model APIs and alignment pipelines evolve, estimated dependency structures should not be viewed as permanent model properties." An entanglement graph is a snapshot of a set of endpoints on a date, not a property of the weights.
Not stated: the roster spans roughly three model generations (Llama-2-70b and Qwen1.5 alongside GPT-5, Claude 4.6 Sonnet and Gemini-3.1-Pro), and the BEI graph's entire top tier is the oldest, weakest tier. Difficulty conditioning is supposed to handle capability, but d_t is estimated from the other 16 models, so for a pair of uniformly weak models the fitted p_m(d) is high everywhere and the residual has little room to be anything but positive. Zhu's release-date confound — a benchmark score matrix's leading factor tracking release date at R² = 0.505 — is the same hazard on the aggregate axis, and nobody has checked whether the failure-manifold version survives restriction to a single release window.
Connections#
- LLM-Judge Validation — the panel-side measurement this complements. Yang's intra-class ρ = 0.66–0.97 is read off a vote matrix and converts jury size into predicted jury accuracy; it cannot say what the jurors agree about. BEI/CIG answer a different question with a stricter null and a ground-truth requirement, and land on that page's own axis from the far side: the judge's over-endorsement of a specific target tracks its entanglement with that target, so a judge's discrimination is a property of the judge–target pair, not of the judge
- Weak-Verifier Ensembling — the direct downstream application: an entanglement-penalised ensemble weight, and the corpus's first attempt at the "price the dependence" exercise that page's open question asks for. The answer so far is discouraging for the premise rather than for the method — the penalty adds 0.1pp over competence weighting alone. Also the home of Sunkavalli's estimator, whose identification wall this instrument gets around by requiring ground truth instead of anchors
- Reference-Free Judge Over-Crediting — the mechanism this measures the population form of. Zhou's Proposition 2 says a shared latent plausibility axis makes every monotone aggregation rule collapse to a threshold on that axis; BEI and CIG measure how much structure the population's errors actually share, with a null model and FDR control rather than a raw pairwise φ. Both agree that lineage is a proxy, not the variable: Zhou's cross-family judges share the basin, and CIG's cross-family pairs are significant
- Same-Model Review Blindness — the same design variable with a better instrument. Greptile's routing rule ("detect the author, route to the other vendor") assumes family membership is what makes a checker blind; this measures the underlying quantity and finds it crossing vendor lines, so "different vendor" is a cheap proxy for "de-entangled" that a CIG pair like Claude-4.6-Sonnet ↔ GPT-5 shows can fail
- Agent Behavioral Homogeneity — the same property measured with a null model instead of anecdotes. That page's evidence is convergence counts from one lab's experiments on its own models; this is 18 models from six vendors tested against conditional independence with FDR-adjusted q-values, and it is the first entry in the corpus on the benefit side of the heterogeneity ledger: mixing across the open/closed-weight line produces fewer significant entanglement pairs than mixing within a tier
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus — the third n_eff instrument, triangulating Kohli's on a different bank (10 open-weight judges, n_eff = 3.51/10) and a new consequence: correlated votes flip a pooled-vs-paired significance-test conclusion in up to 28% of resampled comparisons. Also the corpus's first causal test of a specific mechanism for frontier dependence (GRPO fine-tuning lightly lowers it, doesn't explain it) and the source of the CorrFilter/Bias-Cluster regime-matched filters developed on Weak-Verifier Ensembling
- LLM-as-a-Judge — the practice both papers constrain. Kohli's result is the sharpest statement of it in the corpus: on all three NLI sets and on RewardBench the best single judge matches or beats the nine-judge panel, six of nine leave-one-out removals raise panel accuracy, and the panel's 90.9% accuracy on its own unanimous items is a 9.1% error rate where independence predicts ~0.02% — so "use a panel" and "trust unanimity" are both purchases of redundancy rather than of reliability
- Benchmark Score Redundancy — the aggregate-score version of the same fact. If models' item-level failures are entangled, a models × benchmarks score matrix is low-rank for reasons that start at the item level; and Zhu's release-date confound is the hazard this page's vintage-mixed roster carries unexamined
- Multi-Agent Collective Intelligence — the general mechanism in formal dress: width averages only the per-agent independent term, so population-common structure is a floor no size lowers. BEI is that floor's excess co-failure, measured
- Item Response Theory for LLM Benchmarks — where the dependence this page measures goes unmodelled: an IRT bank calibrated on thousands of community fine-tunes of a handful of bases treats them as independent examinees, so every difficulty and discrimination parameter is defined relative to an entangled population, and a frontier-only population would re-scale it
Open Questions#
- The paper reports three judges and its Spearman degrees of freedom imply two (n = 34 = 2 × 17 on MMLU-Pro, n = 26 = 2 × 13 on MATH-500; three judges would give 51 and p ≈ 1.4 × 10⁻⁴ rather than the reported 0.002). Which judge is missing from the headline association, and does it come back in at the same sign? Settleable by one sentence from the authors or by a per-judge breakdown of ΔPrec.
- Does the entanglement penalty buy anything once the verifier pool is large enough for redundancy to matter? The +0.1pp over competence-only reweighting is measured at
M_J = 3, whereR_maverages over two other verifiers and the exponential penalty has almost nothing to discriminate. The falsifiable version: rerun the reweighting atM_J= 5, 7, 9 and report the entanglement-minus-competence delta as a function of pool size — a flat line kills the method, a rising one is the paper's real result. - Is the BEI graph measuring entanglement or vintage? Its top 10 pairs are entirely Llama and Qwen models from 2023–2024 while no frontier pair appears at all, and difficulty is estimated from the other 16 models, so a pair of uniformly weak models has a high fitted failure probability everywhere. Restricting the roster to one release window (or residualising the error matrix on release date, as Zhu does on the score matrix) would say whether the intra-family signal survives. Partially answered (2026-09-22) by Kohli, from the side rather than head-on. His nine judges are all current frontier endpoints — no three-generation spread — and the panel still shows φ̄ = 0.391 with the three most correlated pairs all cross-family (Claude × Gemini 0.603, GPT-4o × Claude 0.588, Mistral × DeepSeek 0.564) and the same-family excess only +0.047. So on a single-release-window roster the dependence does not vanish and does not organise by lineage, which is evidence against the vintage reading of BEI's intra-family top tier. It is not the experiment this bullet asks for: different instrument (raw verdict-level φ, not difficulty-conditioned excess co-failure), different task (NLI classification, not MMLU-Pro), and no open-weight tier in the panel at all — so it says the cross-family finding survives a vintage restriction, not that BEI's intra-family ranking would. Corroborated again (2026-09-25) by Hossain, Yousefi & Lim, on a third instrument and a fourth judge roster: GPT-5.6-sol, Claude Opus 5 and Grok 4.5 — all current frontier endpoints, no generational spread — show cross-provider mean error correlation (0.42) statistically indistinguishable from the within-provider Gemini bank's (0.40), and the gap survives adjustment for shared item difficulty (0.465 → 0.467). Three independent groups, three instruments (BEI/CIG's difficulty-conditioned excess co-failure, Kohli's Kish n_eff, this paper's Ledoit-Wolf ρ̄), three judge rosters, one direction: vendor boundaries do not predict lower error correlation among frontier judges. The head-on version — rerun BEI itself on Kuai's roster restricted to one release window — is still unrun.
Sources#
- A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges — Chenchen Kuai, Jiwan Jiang (equal contribution), Zihao Zhu, Hao Wang, Keshu Wu, Zihao Li, Yunlong Zhang, Chenxi Liu, Zhengzhong Tu, Zhiwen Fan & Yang Zhou (corresponding), A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges, arXiv 2604.07650 v2 (2026-08-09), 19pp, accepted at COLM 2026; Texas A&M University (1), Marquette University (2), University of Utah (3).
empirical, tier kept. Cited here for §3.1 (the leave-pair-out difficulty estimator, the conditional-independence null, the easiness-weighted BEI and Proposition 1), §3.2 (distractor attractiveness with add-one smoothing, the directional null, surprisal-weighted CIG), §3.3 and §B.4 (the reweighting construction and its calibration/evaluation split), §4.1 (roster, the two disjoint 1,000-question MMLU-Pro subsets, the three judges, the 5-shot temperature-0 protocol), §4.2 (Table 1, Table 2, the entanglement-graph reading and the causal disclaimer), §4.3 (Table 3), §5 (limitations), §B.1 (the 18-model list), and Appendix C (Tables C.1–C.4). - Evidence handling.
empiricalis correct: measured error matrices over real API endpoints, a stated null, Monte Carlo inference with Benjamini–Hochberg FDR control, calibration diagnostics for the nuisance model, and a calibration/held-out split for the downstream method. Two scope limits travel with the tier. The association half rests on ~34 (MMLU-Pro) and ~26 (MATH-500) observations, a figure the paper never reports and that this page back-solved from the (ρ, p) pairs — every claimed contrast between BEI/CIG and the baseline metrics sits inside overlapping confidence intervals at that n, and only the sign, the significance and the cross-benchmark transfer survive. The mitigation half's headline is against the wrong baseline: +3.5pp accuracy over majority voting is +0.1pp over the accuracy-calibrated reweighting reported one row above it, on 500 held-out questions. COI: none identifiable — an academic group with no vendor relationship, auditing 18 third-party endpoints including four vendors' frontier models; no funding statement appears in the captured version. - Parse note. PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 19pp, 7 tables, 4 pictures, formula engine mlx, rapidocr). The ingest worker's
verify.pyreportedwarnand stood down before confirming table reconciliation, so all seven tables were re-reconciled at compile time againstpdftotext -layouton: Tables 1, 2, 3, C.1, C.2 (10 rows), C.3 (9 rows) and C.4 (2 rows) match cell-for-cell with no dropped rows, no shift and no weld. The one flagged cell, Table C.4's7.6 × 10 - 5, is a scientific-notation value split the same way by both parsers, not a merged cell. Two PDF-level artefacts survive into both parses and are not docling damage:Llama-3 1-70Bfor Llama-3.1-70B andDeepSeek-chat-v2 5for v2.5 (lost decimal points). NoAI→AlOCR misreads (the onlyAltoken is the surname "Al-Dahle" in a reference). Figure 2 (the two entanglement graphs) was read at compile time and its 18 nodes confirmed against the Appendix B.1 roster arithmetically (4 GPT + 3 Claude + 3 Gemini + 4 Llama + 3 Qwen + 1 DeepSeek); Figure 3's weight heatmap was read for the competence-dominance observation. Citation slip worth knowing: Gemini-3.1-Pro is cited to Team et al. 2023 and GPT-5 to a 2025 system card, so the reference list does not identify the actual endpoints used. - Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus — Elias Hossain, Niloofar Yousefi & Ser-Nam Lim, University of Central Florida, Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus, arXiv 2609.22512 (2026-09-18), 50pp, no venue stated.
empirical, tier kept. Cited here for §2.1 (the Ledoit-Wolf-shrunk error-correlation matrix, n_eff and eigen-effective rank), §3 (the ten-judge open-weight bank and RewardBench/UltraFeedback/PKU-SafeRLHF calibration sets), §4.1 and Table 9 (the headline ρ̄/n_eff and the pooled-vs-pair-based significance test, Figure 2), §4.2 and Tables 1, 7, 8 (frontier-bank correlation, the cross-provider-vs-within-provider decomposition, the 7.7× co-failure lift, the difficulty-adjustment check), §4.3 and Table 2 (the GRPO/DPO forced-choice comparison), App. F.3 (the 73-subbank robustness sweep) and App. M.4 (the label-noise sensitivity check). - Evidence handling.
empiricalis right and unusually thorough for a three-author academic paper: real API/local-inference judge votes on three independent preference datasets, a stated Ledoit-Wolf shrinkage estimator rather than a raw sample correlation, item-level (not vote-level) bootstrap confidence intervals throughout, a preregistered diversity-contrast hypothesis that the paper reports failing (model-family and prompt-style contrasts both miss the preregistered Δ ≥ 0.10 threshold on RewardBench), and a stated scope-conditions section naming exactly what does not transfer (App. M.1–M.3: the global/subgroup regimes are controlled interventions with unknown natural prevalence, position bias is the only shared vulnerability evaluated in depth, and numerical ρ̄/n_eff values should not be assumed to transfer to differently-composed banks). COI: none identifiable — an academic group with no stated vendor relationship, evaluating third-party open-weight models it ran locally and three proprietary frontier judges (GPT-5.6-sol, Claude Opus 5, Grok 4.5) it holds no competitive stake in; no funding statement was captured in the ingest. - Parse note. PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 50pp, 30 tables, 9 pictures, layout and table engines
mlx, rapidocr,confidence_grade: excellent). Table 3 (§4.4, the weak-dependence/vulnerable-subgroup filter-precision comparison) is cell-collapsed: its three "Weak dependence" rows (5%/10%/20% reversed labels) merge into one grid row per the compiler-prompt's documented failure mode, so the Supermaj. column reads0.980 0.956.901(a dropped0before.901) and the CorrFilter column reads0.976 0. 0.with948and891stranded on the following rows. Reconciled against Table 15 (§F.6, "Filtering precision under the weak-dependence intervention at matched retention"), which reports the identical experiment with rows and columns transposed and uncollapsed: Naive/Supermajority = 0.980 / 0.956 / 0.901, CorrFilter (clean R) = 0.976 / 0.948 / 0.891 at 5%/10%/20% — matching the collapsed cell's fragments exactly and matching the prose's own paraphrase ("CorrFilter and Bias-Cluster are both 0.4–1.8 points less precise than the supermajority": 5% gives 0.4pt/0.7pt, 10% gives 0.8pt/1.1pt, 20% gives 1.0pt/1.8pt against Table 3's uncollapsed Bias-Cluster column of 0.973/0.945/0.883, so all three columns now check against each other and against the prose). No other table cited above needed reconciliation: Tables 1, 2, 7, 8, 9 and 14 were read directly and their cross-references (e.g. Table 1's Open (all) n_eff = 3.75 vs. Table 9's 3.51 — two different subsets, 400 vs. 1,133 pairs, both stated in the captions) resolve without repair. - Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels — Guneet Kohli (single author, Apple), Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels, arXiv 2605.29800 (2026-05-28), 14pp, 12 tables, 5 figures, no venue stated.
empirical, tier kept. Cited here for §3.1–3.2 (ChaosNLI sampling and the 9-judge / 7-family roster), §3.3 (Kish and eigenvalue n_eff, the bootstrap), §3.4 (the item-aware Condorcet null), §3.5 and §4.3 (the stratified permutation test), §4.1–4.2 (Table 2, Figure 1, the gap and its difficulty decomposition), §4.4 (Table 8, the scaling curve and the 1/φ̄ asymptote), §4.5 (Table 3), §4.6 (Table 4, RewardBench, chain-of-thought), §5.1–5.4 (Tables 9, 10, 5 and 11; the same-family/cross-family comparison and Figure 3), §6 (the n_eff/k < 0.5 rule), Limitations, and Appendices B, E, G, I, J and K. - Evidence handling.
empiricalis right and unusually well supported for a single-author paper: 9 real API endpoints × 4 datasets × 5 conditions, a stated null model, a stratified permutation test that preserves per-judge error rates and the difficulty structure, bootstrap CIs, split-half cross-validation of the Condorcet estimate (overfitting ratios 0.960–1.000), a sample-size convergence check, and a documented tie-breaking robustness check (flipping all 28 gold-label ties moves panel accuracy ≤ 0.7pp and n_eff ≤ 1.3%). Three scope limits travel with the tier. The human comparator is simulated: Table 3's human n_eff of 4.03–5.79, which carries the paper's "human panels are ~2× more independent" claim, is obtained by sampling 10 labels per item from the aggregate ChaosNLI distribution under an exchangeability assumption (the paper's own footnote), not by measuring the 100 real annotators — treat it as a reference line, not a measurement. The bootstrap CI covers items, not judges (stated in §3.3): a different nine-model panel could land elsewhere, and no judge-selection uncertainty is quantified. All four tasks are classification or binary preference; the authors name open-ended generation and code review as untested, which is precisely where judge panels are deployed. COI: none identifiable in the direction of the finding — Apple ships no model in the panel, and the result is unflattering to every vendor that does, including the writing-assistance tool the paper discloses (Claude). - Parse note. PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 14pp, 12 tables, 5 pictures, layout and table engines
mlx, formula enrichment on, rapidocr,confidence_grade: excellent). Ingest reportedokon both verify runs with canary-recall 14/14 on a real (non-degenerate) sample. All 12 tables were compared againstpdftotext -layoutat ingest and 11 matched exactly; Table 3 (cross-dataset, p6) was repaired in place — docling welded the "Task type" and "n_eff (Kish)" row labels and split the MNLI column across them — and the repair is recorded in a> [!note]block in the raw. Every table cited on this page was independently re-reconciled at compile time: Tables 1, 2, 4, 5 (pages 4–8) and Tables 6, 7, 8, 9, 10, 12 (pages 12–14) matchpdftotext -layoutcell-for-cell, and the repaired Table 3 was re-verified against page 6 and against Table 2's own MNLI n_eff. Row-count arithmetic checks out: Table 6 has 11 rows with onen < 5bin omitted by the caption, Table 10's classes sum to 1,000 (476 + 359 + 165), Table 12's three breakdowns each sum to 51, and Figure 1's ten bars sum to 1,000 (290 + 140 + 103 + 108 + 68 + 69 + 65 + 61 + 45 + 51). No en-dash corruption, no welded ranges, noAI→Al. Figures read under the image two-pass rule: Figure 1 (error histogram, bar labels transcribed and summed), Figure 2 (scaling curve — confirms the1/φ̄ = 2.6asymptote line and the min–max band), Figure 3 (the 9×9 phi matrix — the source of the "tight five-vendor block at 0.46–0.60" reading, which the prose does not state), Figure 4 (per-bin Condorcet gap, redundant with Table 6). Figure 5 (sample-size convergence) is fully described by Appendix F prose and was not needed.
Cited by 13
- Weak-Verifier Ensembling×8
Cross Model Error Entanglement — the dependence measured from the other side, with a null model and…
- LLM-Judge Validation×7
Finding C's reporting rule is a good rule with a missing half: it tells you to report a correlation…
- Agent Behavioral Homogeneity×6
Cross Model Error Entanglement — this page's property tested against a null instead of counted: 18…
- Reference-Free Judge Over-Crediting×4
A fourth measurement, on ordinary judging with no optimizer and no manufactured errors, lands on…
- Same-Model Review Blindness×4
Cross Model Error Entanglement — the variable this page's vendor-routing prescription is a proxy…
- Open Questions Backlog×2
Cross Model Error Entanglement ×2 (oldest 7d) — The paper reports three judges and its Spearman…
- Benchmark Score Redundancy
Cross Model Error Entanglement — the redundancy's microstructure, one level below the score matrix.…
- The Price of Mixing Agents, and the Principal Nobody Counted
Postscript, 2026-09-22: the benefit side is no longer empty, and the first coefficient is negative.…
- Item Response Theory for LLM Benchmarks
Cross Model Error Entanglement — the measurement of the assumption this page says nobody tests: IRT…
- LLM-as-a-Judge
The standard hedge against the judge-dependence property above is to stop picking one judge and…
- Evals & Benchmarks
Cross Model Error Entanglement — Three 2026 audits of whether LLM errors are independent, measuring…
- Multi-Agent Collective Intelligence
Cross Model Error Entanglement — Proposition 1's non-averaging term, measured on real models…
- Open Questions Dashboard
Cross Model Error Entanglement: Is the BEI graph measuring entanglement or vintage? Its top 10…
Related articles
- Weak-Verifier Ensembling
Weaver (Stanford, 2025): stop training a better verifier and combine the imperfect ones you have — normalize a heteroge…
- LLM-as-a-Judge
Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…
- LLM-Judge Validation
UC Berkeley's 21-judge / 9-provider / ~541K-judgment audit (Norman et al., 2026): LLM-as-a-judge validation is systemat…
- Reference-Free Judge Over-Crediting
Reference answers are a first-order determinant of LLM-judge verdicts: without one, judges systematically over-credit w…
- The Price of Mixing Agents, and the Principal Nobody Counted
Joint answer to two #oq/now items about what a population of agents does that no single agent does. (1) No variance-vs-…
