H
Howardism
Plate IIEvals & Benchmarks中文HOWARDISM

Recurring Production Agent Evaluation

She & Lin (CMU/Meta, arXiv 2609.21267, adopter-voice deployment report): 574 runs of a production analytics agent's 519-question benchmark over 52 days, split chronologically 287 calibration/287 held-out, compare random sampling, historical outcome caching, fixed difficulty-stratified subsets, and Rasch/multidimensional-2PL adaptive testing for *recurring* evaluation of one *evolving* system. Multidimensional-2PL adaptive testing wins on score fidelity from k=200 onward (1.03pp MAE at 38.5% of the benchmark, vs Rasch's 1.37pp and random's 1.79pp), and difficulty-stratified fixed subsets win at small budgets — but the team deployed the fixed subsets anyway, for operational simplicity over accuracy, and backed that choice with unrecalibrated transfer to five other agent families and a calibration window shrunk to one day without loss.

Article metadata
Publication details
Published:September 25, 2026
Filed:Concept
Domain:Evals & Benchmarks
Reading:12 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Recurring Production Agent Evaluation

Sources#

Summary#

Every efficient-evaluation paper this wiki carries so far calibrates on responses of distinct models evaluated on a public, static benchmark (tinyBenchmarks/Polo et al. 2024, ATLAS, BenchPress/CollabEval). Yining She (CMU, work done at Meta) and Lei Lin (Meta) (arXiv 2609.21267, 2026-09-18, empirical) instead calibrate on temporally-ordered runs of one evolving production system: a deployed analytics agent serving tens of thousands of monthly active users, re-evaluated on a 519-question, ~3-hour internal benchmark as its models, prompts, tools, and surrounding infrastructure change. They use 574 historical runs collected over 52 days, split chronologically (not randomly) into 287 calibration runs (Day 1–28) and 287 held-out test runs (Day 29–52) — testing whether a method calibrated on the past generalizes forward in time, not merely to unseen items.

They compare four method families: random sampling; historical outcome caching (reuse a question's recent majority verdict instead of re-executing it); fixed representative subsets (à la Polo et al. 2024, selected by difficulty stratification or clustering, scored by a weighted estimator or by gp-IRT reconstruction); and IRT-based adaptive testing (à la Truong et al. 2025's Rasch/Fisher-information procedure, extended here to a multidimensional 2PL with D-optimal item selection). Multidimensional-2PL adaptive testing gives the best score fidelity from k=200 questions onward (1.03pp MAE, 38.5% of the benchmark) and the best ranking fidelity (0.992 Spearman at k=300). Difficulty-stratified fixed subsets win at small budgets and are close behind at every budget. Despite this, the team deployed the fixed subsets, not the statistically better adaptive method — the paper's central contribution is the first-hand account of that tradeoff and the validation work done to justify it.

Two "adaptive testing" curves — don't conflate them#

The paper's headline multi-method comparison (§5.1, Fig. 2, Table 1: caching vs. fixed subsets vs. adaptive testing) uses Rasch adaptive testing as the adaptive representative — its MAE at k=200 is 1.37pp, narrowly behind fixed-subset gp-IRT's 1.40pp until k=200, then ahead. Separately, §5.3.2 (Fig. 5, Table 4) compares Rasch against a multidimensional 2PL adaptive variant and shows it is substantially better everywhere, especially at small budgets (k=100: 1.97pp vs. Rasch's 2.92pp and random's 2.94pp). The abstract's headline number — 1.03pp MAE at k=200 (38.5% of 519) — is the multidimensional-2PL figure from Table 4, not the Rasch figure from Table 1's headline comparison. A reader who quotes "1.03pp" against Table 1's caching/fixed-subset curves is comparing across two different tables built for two different purposes.

Results#

Score fidelity#

  • At k=100 (19.3%): difficulty-stratified gp-IRT leads the non-adaptive-vs-Rasch comparison at 2.65pp, vs. random 2.94pp and Rasch-adaptive 2.92pp. MD-2PL adaptive is already ahead of all of them at 1.97pp, though it is compared separately (see above).
  • At k=200 (38.5%): Rasch-adaptive (1.37pp) overtakes fixed-subset gp-IRT (1.40pp) in the headline comparison; MD-2PL adaptive is well ahead of both at 1.03pp.
  • Historical caching is dominated on MAE at matched execution fraction across its whole sweep (7.36pp at 20.3% mean execution down to 0.81pp at 81.6%) — its curve sits above every other method over their shared range (e.g. 1.54pp at 69.1% execution vs. 0.60–0.93pp for the other three methods at k=350, 67.4%).

Ranking fidelity decouples from score fidelity for caching#

Caching's Spearman/Kendall correlations are much stronger than its MAE alone would predict: 0.905 Spearman at only 20.3% execution, reaching 0.996 at 79.0% (vs. 0.995 for MD-2PL adaptive at k=400, 77.1%). A method can preserve which run is better while being a poor estimator of how much better — the same score/rank split ATLAS makes in the other direction, where 23–31% of models move >10 ability ranks despite 0.96–0.99 Spearman against accuracy. Together the two papers show score fidelity and ranking fidelity are not simply nested in either direction.

The contradiction with Polo et al. (2024): clustering selectors lose to a scalar difficulty ordering#

Difficulty stratification (fit a Rasch model, stratify by difficulty, pick the median item per stratum) beats both clustering-based selectors — historical-response K-means and IRT-feature K-means, the two selectors adapted from Polo et al. (2024) tinyBenchmarks — from k=150 (28.9%) onward, and from k=200 both clustering selectors fall below random sampling under either estimator. This directly contradicts tinyBenchmarks, which found IRT-based (anchor-point) clustering consistently effective for subset selection. The authors offer two untested hypotheses for the reversal: (1) one evolving agent's historical response matrix may carry less distinct-response-pattern diversity than a population of independently trained LLMs, making multidimensional item parameters and their clusters less representative; and (2) their subsets are a far larger fraction of the benchmark (k=100 is 19.3% of 519) than tinyBenchmarks' (100 of ~14,000 MMLU items, <1%) — random sampling itself gets more accurate at large fractions, narrowing the room for smarter selection to beat it.

The deployment decision: simplicity over accuracy#

Despite MD-2PL adaptive testing's superior score and ranking fidelity from k=200 onward, the team deployed difficulty-stratified fixed subsets at k∈{100,200,300,400}, letting users pick their own execution/fidelity tradeoff. The stated reasons are entirely operational, not statistical: fixed subsets reveal the workload before execution and fit an evaluation system built to run questions in parallel (adaptive testing needs sequential, run-specific ability updates); they preserve the same questions across every run, supporting direct question-level comparison and regression diagnosis as the agent evolves; and caching needs continuous monitoring for stale verdicts that fixed subsets don't. This is the paper's load-bearing contribution: most of this wiki's evaluation-efficiency literature validates a method once against a metric; this is a first-hand account of choosing the metric's runner-up for reasons the metric doesn't capture, then spending the rest of the paper validating that choice.

Validating the deployed choice#

Cross-agent transfer#

The difficulty-stratified subset and its Rasch/gp-IRT parameters — fit only on the original agent's calibration data — were applied without recalibration to 299 runs from five additional agent families with different frameworks, execution harnesses, tool sets, and model configurations (Family A–E: 51/146/56/34/12 runs). Pooled, gp-IRT beats random sampling at 11 of 12 budgets; at k=200, random/weighted/gp-IRT are 1.86/1.49/1.47pp, with gp-IRT close to its 1.40pp on the original held-out runs. The advantage is not uniform: e.g. Family D's gp-IRT and weighted estimators lose to random sampling at k=400 (1.13/1.34pp vs. random's 0.94pp). The authors read the pooled result as evidence the IRT-fitted difficulty scale captures something intrinsic to the questions that generalizes across agent systems, not just across time within one system.

Calibration-window sensitivity#

Holding the calibration end fixed at Day 28 and sliding the start from Day 1 to Day 28 in 3-day steps yields ten nested windows of 287 down to 14 runs; for each, Rasch and gp-IRT are refit and the subset rebuilt. MAE is not monotonic in window length — at k=200, gp-IRT MAE ranges 1.35–1.59pp across all ten windows, and the shortest window (14 runs, effectively one day of history) reaches 1.41pp, statistically indistinguishable from the full four-week window's 1.40pp. The authors' own caveat matters: they attribute this to the agent being at a "relatively mature development stage" with limited day-to-day change during the study, and do not claim it generalizes to an early-stage or fast-churning agent.

Practical recommendations (from the paper)#

  • Recalibrate after material changes to models, prompts, tools, or the execution system, or after a sustained shift in question-level outcomes — not on a fixed calendar cadence.
  • Periodically run the full benchmark and compare its pass rate against the fixed-subset estimate, to measure the drift the subset accumulates during ongoing monitoring.
  • Weigh operational properties alongside statistical fidelity: caching needs recent-outcome monitoring for staleness; adaptive testing needs sequential orchestration and run-specific state; fixed subsets need neither, at the cost of not being the most accurate option.

Scope and limitations#

  • Single organization, single production benchmark. The cross-agent-transfer and calibration-window validations (§6) were run only on the deployed difficulty-stratified fixed subset — historical caching and adaptive testing were never checked for transfer or window sensitivity, so nothing here shows whether adaptive testing's superior MAE, or caching's rank-fidelity-despite-MAE, also transfer or survive a short calibration window.
  • Pass-rate MAE and aggregate rank correlation don't directly measure regression detection or release-gate decisions at an operational threshold; the paper names this as future work.
  • No benchmark questions, outcome matrix, or implementation are released — the methods are fully specified in §3 and reproducible on any benchmark with recorded per-item outcomes, but nothing here is independently checkable against the original data.

Connections#

  • Item Response Theory for LLM Benchmarks — shares the Rasch/2PL/Fisher-information machinery (ATLAS also fits an IRT model, selects items by Fisher information, and stops on an ability-precision target), but the object under calibration differs. ATLAS's item bank is fit once on a static leaderboard population, and its largest measured degradation is a ~1-year temporal holdout across distinct models; this page's ten calibration windows are all drawn from one evolving system's own run history, and its headline surprise — a 1-day window ties a 4-week window — is evidence for a narrower claim than "IRT item banks resist staleness": that this agent's difficulty ordering didn't need much lookback during a mature development phase. Relevant to that page's open question on recalibration cadence, without resolving it — see the annotation there.
  • Adaptive Stopping in Evaluation Sampling — a fourth axis in that page's "routes to a cheaper evaluation" comparison. optstop stops sampling epochs per item; ATLAS stops administering items per model by ability precision; this paper runs both non-adaptive (caching, fixed subsets) and adaptive (Rasch/MD-2PL) methods against each other and picks the non-adaptive option for an operational reason — no per-run sequential state — that none of the three routes on that page's table carries, since all three there are non-adaptive-administration methods and none reports an operational-simplicity tradeoff against a statistically superior alternative.
  • Benchmark Score Redundancy — this page's fixed-subset method is Polo et al. (2024) tinyBenchmarks' method, redeployed on one evolving production agent instead of a cross-model public leaderboard — and it contradicts tinyBenchmarks' clustering-selector result rather than replicating it (clustering underperforms difficulty stratification and even random sampling here; see above). A third data point for that page's item-level-dimensionality debate (CollabEval: item-level matrices need ~16 components; a unidimensional 3PL fits ATLAS's item banks well): here, cluster-based selection on one system's homogeneous response history actively underperforms a one-dimensional difficulty ordering, consistent with the authors' own hypothesis that a single evolving system's response matrix has less exploitable multidimensional structure than a diverse model population.

Open Questions#

  • Does the calibration-window stability (1-day ≈ 4-week) hold for an agent earlier in development, or one changing faster? The paper explicitly attributes the null result to a "relatively mature development stage" and does not test an early-stage or high-churn agent.
  • Do historical caching and adaptive testing transfer across agent families and tolerate a short calibration window the way the deployed fixed subsets do? Both validations in §6 were run only on the difficulty-stratified fixed subset.
  • Does difficulty stratification's win over clustering-based selection generalize beyond one production agent's response history, or is it specific to the narrower response-pattern diversity the paper itself proposes as the explanation for its contradiction of Polo et al. (2024)?

Sources#

  • Efficient Benchmarking in Production: A Study of an Evolving LLM Agent — Yining She (Carnegie Mellon University, work done while at Meta) & Lei Lin (Meta), Efficient Benchmarking in Production: A Study of an Evolving LLM Agent, arXiv 2609.21267, 2026-09-18, 28pp, empirical, adopter-voice deployment report. Cited for the full paper: §2 problem setup (Eq. 1–5, score and ranking fidelity objectives); §3 method definitions (caching eligibility Eq. 8, weighted/gp-IRT estimators Eq. 12–15, Fisher/D-optimal adaptive selection Eq. 16–17); §4 experimental setup (574 runs, 287/287 chronological split, method configurations); §5.1–5.3 (Figures 2–5, Tables 1 and 4); §6 (cross-agent transfer, Table 5–6, Figure 6; calibration-window sensitivity, Table 12, Figure 7); §7 practical recommendations; Limitations. Parse warning: docling-derived (2.126.0/docling-mlx 0.1.1, confidence_grade: excellent), 12 tables. Tables 1 and 4 — the two cited for every headline number on this page — were checked against the surrounding prose and match exactly; both are clean. Tables 3 and 6–12 carry a flattened sub-header row-shift (e.g. Table 3's Kendall's-Tau panel loses its k=200/38.5 row label onto a stray row with blank cells elsewhere in the row) and were not reconciled; no number from Tables 3 or 6–12 is cited by table row on this page — every figure above traces to Table 1, Table 4, or reconciled prose (§5.1–5.3, §6.1–6.2).
§ end
Cited by 5
Related articles