Sources#
- DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
- Driving the Agent Quality Flywheel from Your Coding Agent- Google Developers Blog
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- How Bridgewater Built an AI Analyst That Does Hours of Expert Research in Minutes
- Introducing System One Models & Jev
- Measuring coding agent misalignment in the wild
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- Sidekick's continual learning loop
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- The price is wrong: AI cost calculation has to consider task completion rates, not just token costs
Summary#
Production-sourced evaluation builds a benchmark from real, de-identified usage of a deployed system rather than from synthetic generation or hand-authored prompts. The argument: a benchmark's whole value is predicting real-world performance, and the most representative tasks are the ones users actually issued. DRACO (Perplexity, 2026) is the worked example — its 100 deep-research tasks are distilled from tens of millions of real Perplexity Deep Research queries — and the method is the paper's central contribution, distinct from the rubric design or the grading protocol.
The method (DRACO)#
- Difficulty-proxied sampling. Start from production traffic, biased toward hard cases: DRACO sampled 1,000 queries that drew subsequent negative sentiment or an explicit thumbs-down — i.e., the queries the deployed system handled worst. This mines failures the system already exhibits, which synthetic generation can't target.
- Privacy-preserving reformulation. An automated LLM pipeline strips PII and reduces ambiguity. Critically, no raw user query is ever exposed to a human analyst — anonymization is a precondition, enforced architecturally, not a cleanup step.
- Augmentation toward difficulty + specification. Real queries are often under-specified; augmentation adds context (persona, output format, sources) and broadens scope (temporal, comparative, geographic) so tasks become well-defined and challenging while still reflecting implicit user intent.
- Filtering for objective / tractable / difficult. Keep only tasks with convergent expert success criteria, bounded scope, and genuine difficulty.
- Human gate. A final in-house expert review for security and quality. The pipeline is automatable end-to-end but deliberately keeps a human as the last safety/quality gate.
The payoff DRACO claims: a benchmark that is representative (mirrors the real domain mix and real failure modes) and refreshable (because both research needs and usage evolve, the pipeline can regenerate fresh tasks rather than ossifying).
The core tradeoff: representativeness vs. over-specification#
Production-sourcing buys representativeness but the augmentation step that makes raw queries evaluable also threatens it. The paper is candid: systematic augmentation "reduces ambiguity and improves reproducibility, but it also risks over-specifying tasks and dampening the natural variability of user queries." The de-identification + augmentation pipeline turns a messy, personal, ambiguous query into a clean, bounded, comparable task — and some of what's stripped (ambiguity, personal context, the actual phrasing) is also part of what makes real usage real. Production-sourced is more representative than synthetic, but it is not raw production.
And representativeness of the tasks is orthogonal to validity of the grading: a benchmark can mine exactly the right production queries and still report an untrustworthy verdict if the judge scoring them hasn't been validated. Norman et al. (2026) is the other half — chance-correction, position-swap, and the consistency–bias audit on the grader — so a fully trustworthy production-sourced eval needs both a representative task distribution and a validated judge.
Why production traffic is a moat-grade eval asset#
The method only works if you have a large-scale deployed system generating the traffic — which is exactly the proprietary-data position described in Compounding Data Moat. Real usage at scale is "time-locked, context-specific, and impossible for a copycat to recreate"; here that same asset doubles as an evaluation substrate. A vendor with production traffic can build representative, difficulty-targeted, continuously-refreshed benchmarks that a competitor without deployment simply cannot — and can do so on tasks where its own product currently fails (the thumbs-down sampling). This is the data flywheel pointed at measurement: usage → failure signal → benchmark → product improvement.
The flip side is a credibility question (see DRACO Benchmark): a benchmark sourced from one vendor's traffic, on which that vendor's product wins, carries an obvious incentive — the human gate and the expert rubrics are partly there to answer it.
The product-loop form: Google's flywheel#
Google's Agent Quality Flywheel operationalizes the same principle as a continuous product loop rather than a benchmark. Agents emit OTel traces; each production session "is a genuine request… and each failure is a ready-made test case for the next cycle." Complete traces skip inference and are graded in place; Online Monitors score live traffic continuously, and drifting scores hand failing traces to the eval-fix loop. Google states the ordering explicitly: synthetic scenarios (its User Simulator) are a cold-start bootstrap — "synthetic scenarios get you moving; production data is what makes the loop sharp." That makes three independent arrivals at production-as-eval-substrate: DRACO (capability benchmark), Deployment Simulation (safety forecasting), and the flywheel (continuous quality monitoring) — the method crossing from benchmark construction into day-to-day product tooling.
The buyer-side instance: Databricks builds its own (2026)#
The method arriving from the fourth direction — not a benchmark vendor, not a product loop, but a customer building an eval to decide what to buy. Databricks reports, via The Register (2026-07-13, case-study, secondary reporting), an internal coding benchmark devised from real engineering tasks its own staff performed against its multi-million-line codebase. The stated motivation is contamination-by-optimization rather than representativeness: CTO Matei Zaharia says the company ran the evaluation because models are tuned to existing benchmarks like SWE-Bench, which the article notes OpenAI has called "broken." The eval's output was a purchasing decision — per-task cost and success rate by model and by harness (Cost-per-Task Over Cost-per-Token, Orchestration Sets Token Economics) — which is exactly the tie-breaker role the vendor advice above assigns to it.
Two things separate it from DRACO's pipeline. It is production-sourced tasks, not production traffic: staff engineering work, not de-identified user queries, so no privacy pipeline is needed and no thumbs-down difficulty proxy is available. And its representativeness claim is stated at company scope, not domain scope — Zaharia concedes the results reflect Databricks' own codebase while arguing other companies can run the same evaluation against theirs. That is the honest form of the moat argument on this page: the asset is not transferable, but the method is, and a buyer with a large codebase already owns the substrate.
One step further: production failures as training signal (Shopify, August 2026)#
Every instance above sources test data from production. Shopify's Sidekick account (McNamara & Mazza-Anthony, Shopify Engineering, 2026-08-05, case-study — first-party, unreplicated, self-reported) points the same pipe at the training set: anonymized production traffic is mined daily for hard negatives, a frontier critic panel proposes a repair, the conversation is replayed with the repair injected, and if the judge re-scores it as passing, the replay becomes a reinforcement-learning trajectory with the judge's score as its reward. Failures the critics cannot repair escalate to human annotators scoring against the same rubric. The account's own framing of why: "in a traditional workflow, each failure becomes a bug report or Slack thread. In the flywheel, these failures automatically enter a self-healing pipeline." Full mechanism on Agent Quality Flywheel.
The difficulty proxy is the interesting difference, and it runs the wrong way. DRACO's proxy is a human signal independent of the system being improved — subsequent negative sentiment or an explicit thumbs-down. Shopify's is the judge's own low score: "conversations the judge correctly scores low and that expose where the model is weakest." That is far cheaper and far denser (it applies to all traffic, not the small fraction that draws a complaint) and it is not independent of the thing being optimized — the same instrument selects the training data, gates the repair, supplies the RL reward, and grades the result. A judge blind spot is therefore invisible on both sides of the loop: cases it scores wrongly high never enter the corpus, and the improvement it reports is measured in its own units.
Which instantiates this page's own open question rather than answering it. The thumbs-down bullet below asks whether difficulty-by-current-failure "makes the benchmark a moving target that flatters the next model trained on those failures." Shopify is the corpus's first source that actually trains on the sampled failures, on a daily cadence, with no held-out arm and no clean split anywhere in the account — production-sourced training and production-sourced evaluation drawn from one pipe and graded by one instrument. Nothing is measured, so nothing is settled; what changes is that the hazard now has a deployment rather than a hypothesis. The transferable discipline: once production failures feed weights as well as evals, the difficulty proxy has to come from outside the model's own grader, or the split that would detect over-fitting to the proxy does not exist.
The user-triggered form: Bridgewater's Teach button (2026)#
Every instance above mines production traffic for the team. Bridgewater's PAT (Bridgewater Associates) adds a variant where the expert user decides what counts as a failure, and the interesting part is what the loop does with that judgment.
Two paths run in parallel. The autonomous one matches Google's flywheel: background agents continuously scan completed investor conversations "figuring out where PAT went wrong," producing human-audited benchmarks that feed changes to context repositories and to the harness. The explicit one is a Teach button in the UI, and its trigger condition is the notable design choice — Michael Ran demonstrates it on an interaction where nothing was wrong: the user simply asked for a different set of visualizations, an angle they judged PAT should have proposed unprompted. The signal being captured is a preference the product failed to anticipate, not an error it committed — a class of failure no automatic grader defines and no trace inspection surfaces, because every step succeeded.
Pressing it spawns an agent that reads the conversation for three named categories: behavioral mistakes, context gaps, and user steering that can be front-run. The user edits or accepts the agent's reading, and on submit the back end runs a sequence worth stating in order:
- Write a benchmark that is expected to fail — reproducing the poor behavior first, so the fix has a witness.
- Iterate on the context repositories or the harness itself until that benchmark passes.
- Re-run the rest of the suite to confirm the fix caused no regression.
- Post a Slack message with a pull request containing the proposed changes to PAT.
Step 1 is the one that separates this from the other instances on this page: a red test before the fix converts "the agent learned something" from an assertion into a reproducible artifact, and it is the discipline that makes step 3 meaningful. Steps 2–4 are Agent-Authored Harness Optimization — an agent editing its own harness — with the human gate relocated to PR review rather than to benchmark authorship. The claimed payoff is fleet-wide: "the next time a human comes to PAT with a similar question, we expect them to get the better version of PAT right out of the box," one investor's correction compounding for all of them.
What is missing is any measurement of the loop. No acceptance rate for the generated PRs, no benchmark-suite size or growth rate, no report of what fraction of Teach presses produce a durable improvement, and no account of what happens when two investors' preferences conflict — which, for a tool whose failures are preferences rather than errors, is the failure mode the design most invites. case-study, first-party, unmeasured.
Contrast with the alternatives#
- Synthetic generation (DeepResearchEval, ReportBench, DRBench) — scalable, no privacy exposure, but tasks are model-imagined and may miss real failure modes. (The one instance in this corpus that measured the gap rather than asserting it is OmniVChat, in the section above: a fully generated benchmark plus a 360-recording probe the generator never touched. The deltas matched; the rankings did not.)
- Auto-generated from a live corpus (DeepScholar-Bench) — DRACO's Table 1 files this under synthetic, and on the human-authorship axis that is right; but its tasks are derived from recent arXiv papers on a monthly refresh restricted to post-training-cutoff publications, so it is neither model-imagined nor contaminatable. It buys refreshability the way this page's method does — from a stream that renews itself — without needing production traffic or a privacy pipeline. The cost is that the domain is whatever the corpus is about, and nothing about real user intent survives.
- Hand-authored from interviews/searches (xBench, ResearcherBench, DEER) — human-authored and realistic, but bounded by author imagination and not drawn from a live production system.
- Production-sourced (DRACO) — the only one of the three that mines the actual distribution and the actual failures, at the cost of needing deployment access and a privacy pipeline.
The synthetic pole, taken all the way, with a recorded probe bolted on (September 2026)#
This page's "Synthetic generation" bullet has always carried the same unpaid debt: tasks are model-imagined and may miss real failure modes — asserted, never measured, because nobody generating a benchmark also builds the recorded counterpart to check it against. OmniVChat (He, Chu, Chen et al., Alibaba Qwen Team + CUHK + SJTU, arXiv 2609.21465, 2026-09-18, empirical) is the first source in this corpus that does both.
Why production sourcing was unavailable here, which is the interesting part. The task is native audio-visual dialogue — a user's phone camera and microphone, query embedded in both streams. The paper's stated constraint: "Open-source recordings of people using their own devices are scarce," with a long tail of noise, camera motion and device posture. This is the case this page's method structurally cannot serve: the deployed systems that do hold this traffic (live-voice products) do not release it, and no academic lab has a deployment. So sourcing collapses to generate-or-record, and the paper does both and compares them.
The generated side. OmniVChat-Studio, a four-agent engine — Director (all text I/O), Renderer (prompt → synchronized audio-visual clip), Reviewer (watches the rendered clip, writes a caption and a seven-part quality report, answers focused questions), and a deterministic Validator — running a 14-stage single-turn pipeline and a 14-stage multi-turn one, with bounded repair budgets per route and a discard-the-run rule when a budget is exhausted. What is worth lifting from it independent of the modality:
- The distribution is declared, not observed. Each subcategory carries a configuration of attributes with target probabilities, and sampling is a Gumbel-top-k race over the full joint assignment space with soft preference multipliers and hard exclusions (Eq. 7). This is the exact inverse of DRACO's difficulty-proxied sampling from real traffic: representativeness is specified rather than estimated, so it is auditable and wrong in a different way — the paper concedes "a small batch may miss rare combinations" and that "later failures and quality filters can also change which assignments reach the final delivered set."
- Observation is separated from judgment, by construction. The Reviewer reports what the rendered clip contains; the Director compares that report against the plan and decides. The design note is explicit that "the reply follows the rendered evidence instead of the intended script" — reference answers are grounded in what actually got rendered, not in what was asked for. This is the synthetic-data analogue of the labelling discipline this page demands of a production pipeline, and it is the part most transferable to a text-only setting.
- The gates are real and they discard. Deterministic validation before semantic review; each repair returns to the stage that caused the defect; the engine "discards an incomplete run instead of saving a partial instance." Then two post-hoc passes: a deterministic audio tail-artifact repair (695 clips across 571 of the 2,800 delivered instances; median transient score 86.6 → 0.02) and a human inspection of every delivered instance, justified in the paper's own words because the semantic gates "use the Reviewer's description and the Director's judgment" and "these agents can make the same error."
That last sentence is the honest statement of the failure mode this page's synthetic bullet gestures at: a generated benchmark's quality gate is correlated with its generator. The remedy applied is a human pass on 2,800 clips — cheaper than a privacy pipeline, and the reason this route scales where production sourcing cannot.
The recorded side, which is what makes this citable. OmniVChat-Bench-Human: 360 phone recordings, 12 single-turn subcategories × 30, all Chinese, performers improvising from a plain-language actor guide with no clip-level script — "unlike synthesized dialogues, the records have no timeline, scene setup, or configuration block." Annotation is three independent captions merged, then human-corrected against the recording, then a reference reply and tiered rubric written from the corrected caption without the video. One number from that pipeline belongs on this page: before human correction, one caption omitted the user's speech in 62 of 360 clips (17.2%) — an annotation defect rate on real recordings that no synthetic pipeline has to pay, and a concrete reason recorded data is not automatically the higher-quality substrate.
The comparison, and its limits. Training on 5,600 engine-generated dialogues moves the synthetic benchmark 0.465 → 0.652 and the recorded probe 0.402 → 0.632; against a language-matched control (the 1,034 Chinese synthetic instances) the two curves are near-identical, 0.437 → 0.634 vs 0.402 → 0.632. On the transfer question this page cares about, that is a genuine result: a benchmark built entirely from generated stimuli predicted improvement on recorded stimuli of the same subcategories. Three things bound it, two of which the paper states itself. The probe is 360 instances covering 12 of 17 subcategories, single-turn only, one language, no latency and no live interruption. Synthetic and recorded halves share the same rubric form and the same grader, which the paper notes "does not make the tasks independent or guarantee that grader bias cancels." And the ordering evidence runs the other way from the level evidence: Gemini-3.5-Flash leads the synthetic benchmark while Gemini-3.7-Flash leads the recorded probe, so the two instruments agree on whether training helped and disagree on which model is better. Transfer of a delta is not transfer of a ranking, and this is the corpus's first source with the data to tell them apart.
Connections#
-
Evaluation-Time Answer Leakage — the leakage channel this page's method does not close, and the one case where freshness is no defence at all. Production-sourced refresh prevents contamination because a brand-new task is hard to pre-memorize; but a task built from a recent upstream commit is maximally exposed to run-time retrieval, since the commit, its diff and its tests are all still upstream and reachable from the sandbox. On SWE-Bench Pro that cost six of seven models 14–26 points. The pairing to keep: refresh defends the training boundary, isolation defends the run boundary, and an agentic benchmark needs both
-
DRACO Benchmark — the worked example; this method is its central contribution
-
GDPval Benchmark — the third sourcing route, adjacent to this page's and distinct from both of its usual alternatives. GDPval is neither production-sampled nor imagined: OpenAI recruited practising professionals (4-year floor, 14-year mean, fewer than 10% of applicants accepted) and had each build a task around work product they had actually produced on the job, then put all 1,320 through model screening and an average of five expert reviews. What that buys is what production sampling cannot: representativeness against an external frame (O*NET work activities, occupations weighted by wage mass), a wage-denominated dollar value per task ($398 mean), and a human reference deliverable to grade against — production traffic has queries but no gold answer. What it costs is the trade this page's tradeoff section names, inverted: the distribution is chosen by the benchmark's authors rather than observed, so representativeness is argued (via the Acemoglu–Autor validation of the digital-task classifier) rather than sampled. Sourcing from the expert rather than from the traffic also moves the privacy problem — no PII pipeline is needed, but the tasks must be scrubbed of details identifying the expert who wrote them
-
Automated Failure Attribution — the opposite trade, taken deliberately and worth reading as the counterweight to this page. WHO&WHEN PRO reaches 12,326 failure traces with golden agent/step/mode labels precisely by synthesizing them — one error injected into a warm-started run that had already succeeded, so the decisive step is correct by construction rather than by annotation. Nothing production-sourced can match that label fidelity at that scale; nothing synthesized can speak to what actually fails in deployment or how often. The two failure modes are exact mirrors: a production-sourced eval has real tasks and contestable labels, a synthesized one has perfect labels and a designed distribution
-
Cost-per-Task Over Cost-per-Token — the same method as vendor advice: Anthropic's model-selection guidance concedes that public benchmarks saturate at the Opus/Fable tier and tells buyers to curate the deciding eval from their own production traffic instead — production-sourced evaluation as the tie-breaker for a purchasing decision, not just a benchmark-construction technique. Databricks is that advice executed by a buyer, and the numbers it produced (per-task cost and success rate by model) live on that page
-
Orchestration Sets Token Economics — the other half of what a buyer-built eval measured: with the tasks drawn from its own codebase, Databricks could vary the harness as well as the model, which no public coding benchmark exposes. Sourcing the tasks from real engineering work is what makes a cross-harness cost comparison mean anything
-
Interactivity Benchmarks — where the generated-stimulus benchmark above is catalogued as an instrument, with its judge study, its checkpoint-noise result and the COI of one lab owning every piece
-
LLM-as-a-Judge — the grading half of the pipeline; production-sourced tasks + rubric-judge grading make an automatable (human-gated) eval
-
Deep Research Agents — the system class whose production traffic DRACO mines
-
Compounding Data Moat — production usage as a time-locked proprietary asset; this is that asset repurposed as an evaluation substrate
-
Evals as Product Spec — "build your measurement framework before launch / from real usage"; production-sourced evaluation is that principle at benchmark scale
-
Task Time-Horizon Scaling — sibling concern: benchmarks saturate, so the ability to refresh from live usage is what keeps an eval alive
-
Automated Behavioral Audit — the alignment-side analog notes its synthetic scenarios "may not match real-traffic distributions" — exactly the gap production-sourcing closes
-
Telemetry vs. Survey Measurement — Faros AI's telemetry-over-survey stance is the engineering-metrics sibling: measure from the real system, not from self-report
-
Deployment Simulation — the alignment-side application of the same method: OpenAI replays de-identified production conversations to forecast safety behavior pre-release, where DRACO replays them to build a capability benchmark; same PII pipeline, same proprietary-traffic moat
-
Conversation-to-Delegation Shift — its measurement-obsolescence argument is the same instinct one step further: as usage becomes delegation, even which metrics to read (complexity, runtime, concurrency, output) must be re-sourced from real agentic behavior, not interaction counts
-
Agentic Code Generation as Compilation — the reproducibility argument for why an eval substrate needs a deterministic pipeline underneath it: Bridgewater buys two-agent code identity explicitly so that "as we're scaling and hill climbing and evaluating, we have something much more dependable than vibes-based or LLM-as-judge evals." A production-sourced benchmark still needs the run-to-run noise floor low enough that a delta is attributable to the change rather than to sampling
-
Agent Quality Flywheel — the continuous product-loop form: OTel production traces graded in place, Online Monitors on live traffic, synthetic simulation demoted to cold-start bootstrap. It also carries the version that crosses from test data into training data (Shopify's self-healing pipeline), where the difficulty proxy is the in-loop judge's own score rather than an independent human signal
-
Failures That Look Like Success — the failure class production-scale traces could quantify: silent contract violations that demo-sized evals only sample
-
Misalignment in Production Agent Traffic — the method with its curation stage deleted, pointed at alignment: no sampling design, no augmentation, no human QA gate, just two judge rubrics run over 8,600 real coding-agent transcripts. It is the cheapest possible version and it inherits the liability this page's tradeoff section names in its purest form — the distribution is simply whatever traffic arrived, and one heavy user with unusually strict code-review rules supplies 41 of the 76 charted severe monitor-evasion cases. Also the clearest case that production-sourcing is not free of construct problems: the behaviour it counts is defined against an oversight mechanism the user had to install first
-
Context Advantage, Not Taste — production telemetry as context transfer: sourcing evals from real usage moves what the human knows about users into where the model can read it, spending the human's asymmetry by design
-
LLM-Judge Validation — the orthogonal quality axis: this page fixes which tasks the benchmark contains; judge validation fixes whether the grading of them is trustworthy — a representative task graded by an unvalidated judge is still an unreliable eval
-
Measuring Beyond Accuracy Saturation — the sibling answer to "what to do when a benchmark saturates," from the opposite end: this page refreshes the task set from live production usage; Nadgir et al. re-instrument the existing task set along six non-accuracy axes. Both reject the retire-and-replace reflex — new tasks vs new metrics
-
Benchmark Contamination and Decontamination — the prevention vs correction pairing against data contamination: production-sourcing (and dynamic benchmarking generally) avoids leakage up front by drawing fresh, hard-to-pre-memorize tasks and refreshing them; UBD instead repairs a model already inflated by exposure to a static benchmark, without a clean reference. Complementary defenses against the same leakage threat
-
How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them? — the cluster synthesis: production-sourced refresh is the new-tasks move of the five-part replacement portfolio (contamination prevention + representativeness), paired there with judge validation as the two halves of a trustworthy eval
-
Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — refresh-from-production as the general form of contamination prevention in a governance setting, and its limiting condition: the pipeline needs deployment traffic at scale, which a regulator drawing a perimeter does not have and the regulated parties do
-
AI-Assisted Error Analysis — the other half of the same problem. This page asks how to sample production traffic into an eval set; that one asks how to read the sample once you have it, using an agent to cluster traces, build a bespoke review UI, and scale each human annotation back across the corpus. Its warning applies to both: LLM-chosen clustering features are a guess, so representativeness of the reviewed subset is assumed rather than shown
-
Skill Lift — the mirror-image sourcing choice, and the sharpest contrast in the corpus. DRACO draws tasks from real traffic independent of the system being scored; NVIDIA SkillEvaluator generates them from the artifact being scored (
create-eval-dataset./my-skill), so the skill fixes both the task distribution and the expected outputs. Perfect topical alignment, zero independence — the exact inverse of this page's representativeness-with-contestable-labels trade, and it converts the vendor's own maxim ("a skill can only be measured as precisely as its evaluation set describes the job") into a ceiling on what the benchmark can show -
Typed Decision Verifiers — a vendor adopting this page's thesis as policy: TypeSafe publishes no public-benchmark scores for Jev and tells users to build their own evals ("System One tasks are much easier to evaluate") — while an entrant-run shared board is the only third-party number the category has
Open Questions#
- How much does augmentation distort the distribution it claims to represent? Is there a measurable representativeness loss between raw queries and augmented tasks?
- Difficulty-by-thumbs-down biases toward current failures — does that make the benchmark a moving target that flatters the next model trained on those failures? Instantiated, not answered (2026-08-13): Sidekick's continual learning loop runs exactly this loop in production — hard negatives sampled by a judge's low score, repaired, replayed, and folded into weights daily via SFT then GRPO with that same judge as the reward — and reports no held-out split, no independent difficulty proxy, and no arm that would separate "the model got better" from "the model got better at the sampler." So the hazard has a deployment now and still no measurement. The falsifiable form sharpens: hold out a slice of production traffic sampled by a signal the training loop never sees (user thumbs-down, or a judge from a different family), and compare the improvement measured there against the improvement measured on the in-loop judge. Half-answered (2026-09-23), and the half that got run came back clean. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue (
empirical) runs the same shape — a training corpus and a benchmark synthesized by one engine from one set of subcategory configurations, optimized with the grading judge as the reward — and it does hold out the distribution slice this question asks for: 360 human phone recordings, improvised by performers from a plain-language guide, produced by no part of the training loop. The two improvements match closely: 0.465 → 0.652 on the in-loop synthetic benchmark, 0.402 → 0.632 on the recorded probe, and 0.437 → 0.634 on a language-matched synthetic control. So on this instance, distributional self-flattery did not happen. The judge half is untouched and the authors say so: the sameqwen3.7-maxgrades reward, synthetic benchmark and recorded probe, and their six-judge sensitivity study "does not remeasure the OmniVChat-RL gain from 0.465 to 0.652 with other judges." The falsifiable residue is now exactly one experiment — re-grade both the in-loop and the held-out slice with a judge from a different family — and it is cheap, since the reply sets already exist. Note also what a matched gain does and does not buy: the two instruments agreed on the delta and disagreed on the ranking of two models from the same vendor family, so a held-out slice validates an improvement claim without validating a leaderboard. - Can the privacy pipeline (no human sees raw queries) be trusted/audited well enough for regulated domains (medicine, law) where the source traffic is most sensitive?
Sources#
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan et al. (19 authors, OpenAI), arXiv 2510.04374 v1, 2025-10-05, 29pp (
empirical). Cited here for the sourcing contrast only: §2.2 (expert recruitment bar and selection rate), §2.3–2.4 (tasks built from the expert's own work product; model screening plus a mean of five human expert reviews per task, minimum three), §2.1 and A.7 (occupation selection by wage mass and the GPT-4o digital-task classifier validated against Acemoglu & Autor 2011), A.4 Tables 3–6 (task value, duration, O*NET coverage), §4 (the open 220-task gold subset scrubbed of expert-identifying detail). Its own results and COI are on GDPval Benchmark - SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents — Zheng, Shang, Jiang, Tian, Zhu, Ma, Yuan & Zhang (ECNU / Shanghai AI Lab / Fudan), SWE-Bench Pro Verified, arXiv 2609.08149, 2026-09-08 (
empirical, 37pp). Cited here for the limit of freshness-as-defence: the four run-time leakage channels and the 103-and-49-tasks-to-zero confirmed-access result. Full treatment on Evaluation-Time Answer Leakage - DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity — §3 (task construction: sampling, pre-processing, augmentation, filtering, curation), §6.1 (generalization limits; augmentation over-specification caveat)
- Driving the Agent Quality Flywheel from Your Coding Agent- Google Developers Blog — "From the inner loop to the production loop": OTel traces as eval input, Online Monitors, synthetic-as-bootstrap (
vendor-claim) - Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias — Norman et al. (arXiv 2606.19544, June 2026,
empirical): the grading-validity half of eval quality, orthogonal to task representativeness; see LLM-Judge Validation - Measuring coding agent misalignment in the wild — The Docent Team, Transluce, 2026-08-04 (
empirical): the method stripped to its minimum — two LLM-judge rubrics run over 8,600 real coding-agent transcripts with no sampling design, augmentation or human QA gate — and the clearest instance of the representativeness liability, where one heavy user supplies 41 of 76 charted severe cases and the counted behaviour is defined against oversight the user had to install. Full treatment on Misalignment in Production Agent Traffic - Sidekick's continual learning loop — Andrew McNamara & Cody Mazza-Anthony, Shopify Engineering, 2026-08-05,
case-study(first-party account of the authors' own production system; no replication, no controlled arm, no held-out split). Cited for the training-side extension only: hard-negative mining by judge score, the critic-panel → arbiter → hinting → replay repair loop, Toloka escalation for unrepairable failures, and the daily SFT+GRPO fold-in. Its cost and quality figures are not used here. Full source treatment on Agent Quality Flywheel - The price is wrong: AI cost calculation has to consider task completion rates, not just token costs — Thomas Claburn, The Register, 2026-07-13 (
case-study, secondary reporting of Databricks' benchmark blog post; the primary is not in the corpus): the buyer-side instance — an internal coding benchmark devised from staff engineering tasks against a multi-million-line codebase, motivated by models being tuned to SWE-Bench, with the results driving model and harness selection - How Bridgewater Built an AI Analyst That Does Hours of Expert Research in Minutes — McManus, Ran & Weight (Bridgewater Associates), LangChain channel, 2026-07-24, 25:44 talk,
case-study. Cited for the Teach-button loop (15:24–18:08) and the autonomous background-agent variant (5:01–6:24). First-party and entirely unmeasured — no PR acceptance rate, no suite size, no durability figure. See Bridgewater Associates - Introducing System One Models & Jev — Diogo Almeida, TypeSafe AI, 2026-09-15,
vendor-claim: FAQ stance against public benchmarks (cited for the Connections entry only)
Cited by 36
- Agent Quality Flywheel×3
Synthetic User Simulator scenarios bootstrapped the whole first cycle. How much of the 21%→5% delta…
- Context Advantage, Not Taste×3
What changed is the durability. In June the frame was preferred because it "gives us a clearer path…
- Deployment Simulation×3
Production Sourced Evaluation — the same "evaluate on real de-identified usage" method, applied to…
- Measuring Beyond Accuracy Saturation×3
Production Sourced Evaluation — the sibling answer to "what to do when benchmarks saturate": that…
- AI-Assisted Error Analysis×2
Production Sourced Evaluation — the sibling question about the same traces: that page
- Benchmark Contamination and Decontamination×2
Two transferable points. The remedy is prevention by construction, not correction — pick a starting…
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them?×2
Concept articles: Benchmark Score Redundancy (Zeng & Papailiopoulos, arXiv 2606.24020), Measuring…
- Bridgewater Associates×2
Production Sourced Evaluation — PAT's learning loop, in both its autonomous form (background agents…
- Cost-per-Task Over Cost-per-Token×2
Anthropic's own guidance says public benchmarks are "helpful directional guides" that break down…
- Deep Research Agents×2
Deep research is a long-horizon, autonomous, multi-step task — exactly the regime Task Time Horizon…
- DRACO Benchmark×2
Production Sourced Evaluation — DRACO's central methodological contribution: tasks built from real…
- Evals as Product Spec×2
Golden sets are not enough, stated as a rule. "That ground truth should include randomly sampled…
- Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It×2
Contamination is the sibling integrity axis and is the one where the naive fix actively misleads:…
- Perplexity×2
Production Sourced Evaluation — DRACO's method: a benchmark built from Perplexity's own production…
- Typed Decision Verifiers×2
On benchmarks. TypeSafe says it deliberately publishes no public-benchmark numbers, only one-off…
- Agent-Authored Harness Optimization
bridgewater pat ai analyst — McManus, Ran & Weight (Bridgewater Associates), LangChain channel,…
- Agentic Code Generation as Compilation
Determinism is being purchased as an eval substrate, not as an end. Weight's stated payoff:…
- Automated Behavioral Audit
Production Sourced Evaluation — the synthetic-scenario caveat noted here ("may not match…
- Automated Failure Attribution
Production Sourced Evaluation — the methodological contrast. This corpus is synthesized by…
- Compounding Data Moat
Production Sourced Evaluation — the same time-locked proprietary-usage asset, repurposed as an…
- Conversation-to-Delegation Shift
This is the same "measure what the system actually did, not the proxy" instinct as Telemetry Vs…
- Evaluation-Time Answer Leakage
Remedy · Fresh tasks (Production Sourced Evaluation) or post-hoc correction (Benchmark…
- Failures That Look Like Success
What fraction of production agent failures are silent-contract violations vs. loud errors? The…
- GDPval Benchmark
Production Sourced Evaluation — the adjacent-but-distinct sourcing method. DRACO mines…
- Interactivity Benchmarks
Production Sourced Evaluation — the sourcing axis OmniVChat-Bench sits at the far end of: every…
- LLM-as-a-Judge
Production Sourced Evaluation — judge protocol pairs with production-sourced tasks to make DRACO an…
- LLM-Judge Validation
Production Sourced Evaluation — the orthogonal axis of eval quality: production-sourcing fixes task…
- Misalignment in Production Agent Traffic
Production Sourced Evaluation — production traffic as an evaluation substrate, pointed at alignment…
- Evals & Benchmarks
Production Sourced Evaluation — Building benchmarks from de-identified real production usage rather…
- Open Questions Backlog
Production Sourced Evaluation ×3 (oldest 106d) — How much does augmentation distort the…
- OpenAI
A measurement asset. Its scale of production traffic is what makes Deployment Simulation work at…
- Orchestration Sets Token Economics
Production Sourced Evaluation — how the production counterpart above was able to compare harnesses…
- Skill Lift
Production Sourced Evaluation — the opposite sourcing choice, and the sharper contrast in the…
- Task Time-Horizon Scaling
Production Sourced Evaluation — the refresh-from-live-usage method that answers this page's open…
- Telemetry vs. Survey Measurement
Production Sourced Evaluation — the same "measure from the real system, not a proxy" instinct…
- Transluce
Production Sourced Evaluation — the method it practises in its most stripped-down form: a judge…
Related articles
- LLM-as-a-Judge
Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…
- LLM-Judge Validation
UC Berkeley's 21-judge / 9-provider / ~541K-judgment audit (Norman et al., 2026): LLM-as-a-judge validation is systemat…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
