Sources#
- CS329A Self-Improving AI Agents — Part 8: Agentic Evaluations and Long-Horizon Tasks
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
Summary#
GDPval is OpenAI's benchmark for whether a model can do real work that people are paid for. Instead of scoring an answer against a key, it puts the model's deliverable next to a deliverable produced by a practising professional in that occupation and asks blinded expert graders which is better — a pairwise win rate, reported as wins-only and as wins-plus-ties. The primary paper is Patwardhan et al., GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (arXiv 2510.04374 v1, 2025-10-05, 19 authors, all OpenAI, empirical).
The wiki has been quoting GDPval's descendants for months: GDPval-AA (the Artificial Analysis Elo board built on it) is cited on Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Kimi (Moonshot AI), The Open-Weight Frontier Gap and Cost-per-Task Over Cost-per-Token, among others. This page is the ancestor those numbers descend from.
The derivative board has been rescored, and the versions are not comparable. Artificial Analysis's original GDPval-AA board (quoted in Opus 4.8's system card) and its v2 successor (quoted in Opus 5's, and in Kimi K3's and The Open-Weight Frontier Gap's 61-Elo gap) put the same model — Opus 4.8 — at 1890 and 1593 respectively. v2 rebuilt the run: 220 tasks from the gold set, shell access and web browsing in an agentic loop, Elo from blind pairwise comparisons. So a GDPval-AA number is only meaningful with its board version attached, and v1 and v2 Elo must never be differenced against each other — the same caution The Open-Weight Frontier Gap applies when it declines to difference an IRT cyber Elo against an Arena gap.
Provenance: rebuilt on the paper, 2026-09-10#
This page was originally compiled (2026-08-17) entirely from Aakanksha Chowdhery's lecture 8 of CS329A — a lecturer reading slides, through a YouTube auto-caption transcript, practitioner-opinion — with several figures flagged as ASR-damaged. The primary paper was ingested as a baseline-lane source on 2026-09-10 specifically to settle those flags, and every number below now comes from the paper unless attributed to the lecture.
The verdict on the lecture is that it held up well. The task, occupation, sector and gold-subset counts, the dollar value, the reference-file share, the specification share, the two endpoint win rates, the quality distribution and the cost/speed ratios were all correct as spoken. Three things did not survive: the sector-selection rule, the O*NET digital filter, and — most consequentially — the reading that a single roughly-linear line runs from 12% to 48%. Each is struck through in place below.
COI. OpenAI built the benchmark, chose the occupations, recruited and paid the experts, defined the metric, ran the graders, and evaluated four of its own models against three competitors'. The automated grader is GPT-5-high. Every "approaching industry experts in deliverable quality" claim originates with the vendor whose models are being measured. Cutting the other way, and unusually: the headline is won by a competitor — Claude Opus 4.1 at 47.6%, 8.8 points ahead of GPT-5 high — which is not what a benchmark built to flatter its author looks like. Both halves are load-bearing; see the Sources note for where the COI does bite (the cost table, which covers OpenAI models only).
The design#
| Paper (arXiv 2510.04374) | As described in lecture | |
|---|---|---|
| Sectors | 9, each contributing over 5% of U.S. GDP by Q2 2024 Value Added (FRED), together ~75.7% — real estate & leasing 13.8%, government 11.3%, manufacturing 10.0%, professional/scientific/technical 8.1%, health care 7.6%, finance & insurance 7.4%, retail trade 6.3%, wholesale trade 5.8%, information 5.4% | |
| Occupations | 44 — the five highest-wage predominantly-digital occupations per sector (four in retail trade), collectively earning $3T annually | 44 ✓ |
| Tasks | 1,320 full set (≥30 per occupation); 220 open-sourced gold subset (5 per occupation) | ~1,320 / ~220 ✓ (ASR rendered the first as "30 20") |
| Task source | The actual work product of the expert who authored the task, classified against O*NET occupational tasks for coverage | ✓ |
| Expert bar | Minimum 4 years in the occupation, 14 years on average; resume screen for promotion and management history, video interview, background check, training and a quiz; fewer than 10% of applicants accepted; ≥5 qualified professionals per occupation | |
| Occupation filter | GPT-4o classifies every O*NET task as digital or non-digital; an occupation qualifies when its weighted digital-task share exceeds 0.60 (weights from O*NET relevance/importance/frequency). Validated against the Acemoglu & Autor (2011) task-content framework | |
| Duration | 7 hours average expert completion time (intro); gold-subset mean 9.49 h, median 5.0, max 100; full-set mean 8.63 h, max 605 h | ~7 hours ✓ |
| Value | Gold subset mean $398.46 (median $174.81, max $4,114.20); full set mean $391.44, max $32,028.70 — time × median OEWS hourly wage | ~$400 ✓ |
| Reference files | 67.7% of tasks require at least one; gold-set mean 1.92 files | ~70% ✓ |
| Specification | 89.07% of gold-set tasks rated well-specified (8.28% under-, 2.66% over-); full set 89.34 / 8.41 / 2.26 | ~89% ✓ |
| O*NET coverage | Gold subset spans 208 of 1,470 O*NET tasks (14.15%), 25 of 35 skills (71.4%), 26 of 41 work activities (63.4%) | not given |
| Quality control | Every one of the 1,320 tasks through model-based screening plus an average of five human expert reviews (minimum three), in three named stages | "multiple rounds" ✓ |
Representative tasks named in lecture: a 3D model of a cable-reel stand for an assembly line (manufacturing engineer); a competitive landscape (financial analyst); assessing images and drafting a consultation note (registered nurse); cutting an intro reel from a script (film/video editor); a response to a dissatisfied customer (customer service); an itinerary for a family of four (concierge); pricing inconsistencies across purchase orders (audit); a property sales brochure (real estate agent); a vendor-fair layout (recreation worker).
Two design choices that are easy to miss. The tasks were sourced by occupation wage mass, not by AI exposure — the selection instrument is "which knowledge work is the economy paying the most for", which makes GDPval's 44-occupation list a different cut of the labor market from the ones on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated and Task Saturation: Broad but Shallow AI Diffusion. And the paper deliberately keeps NSFW and political content where the occupation genuinely involves it (film, literature, law, politics) rather than scrubbing the distribution.
The headline result — and what "roughly linear" actually covers#
Blinded pairwise expert comparison on the 220-task gold subset, wins + ties (Figure 5; three samples per model per prompt × three human graders = nine comparisons per prompt per model):
| Model | Win-or-tie rate vs. industry professional |
|---|---|
| GPT-4o | 12.4% |
| Grok 4 | 24.3% |
| Gemini 2.5 Pro | 25.5% |
| o4-mini high | 27.9% |
| o3 high | 34.1% |
| GPT-5 high | 38.8% |
| Claude Opus 4.1 | 47.6% |
The single most important correction this compile makes: win rate rises roughly linearly from ~12% (GPT-4o, 2024) to ~48% (Claude Opus 4.1, 2025) (superseded 2026-09-10 by GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks). Those are two different charts. Figure 5 is a cross-vendor snapshot sorted by score, and its top entry is a competitor's model. Figure 6 is the time series that carries the linearity claim, and it plots three OpenAI models only — GPT-4o (06/2024, 12.4%), o3 high (~04/2025, 34.1%), GPT-5 high (~08/2025, 38.8%). Claude Opus 4.1 is not on it. So the paper's "improving roughly linearly over time" is a three-point fit through one vendor's release history ending at 38.8%, and the 48% figure is a different quantity: the best score any evaluated model achieved, on a snapshot with no time axis.
That matters because the linear-versus-exponential contrast with METR (Task Time-Horizon Scaling) is the most-cited thing GDPval does in this wiki, and it is being carried by three points. The contrast survives — three points do not describe a doubling — but it should be quoted with its n.
The lecture's framing of the contrast is still the right one, and Chowdhery states it explicitly: "that's a very different trend compared to METR where we were talking about this doubling trend every seven months. It's more of a linear trend roughly compared to the exponential trend that METR was talking about." The two are not the same measurement — one is task duration at fixed reliability, the other output quality against a human at fixed task — and the deflationary reading is deliberate: on work that actually takes hours to weeks, the models are only so useful, and the usefulness is broken down by profession rather than aggregated into a horizon.
An erratum in the paper's own prose. §3.1 says model deliverables "outperformed or matched expert humans' deliverables in just over half the tasks" immediately after quoting 47.6% — which is just under half. The figure is the authoritative number; the sentence overstates it.
The comparison is system-level, not model-level. Footnote 2 is explicit: Claude was sampled through its consumer UI to enable the file-creation-and-analysis feature, while OpenAI models ran via API with web search, a code interpreter, background sampling and a set of preinstalled libraries. Blinding is also imperfect and the paper says so — OpenAI outputs often used em dashes, Claude outputs adopted first-person phrasing, Grok occasionally named itself; filenames were scrubbed but style was not altered. Some part of the 8.8-point gap between Claude Opus 4.1 and GPT-5 high is a harness and surface difference rather than a model difference, and nothing in the paper separates them.
Three gradients the aggregate hides#
Duration (Figure 13, win-or-tie rate by expert completion time):
| Model | 0–2 h | 2–4 h | 4–8 h | 8+ h |
|---|---|---|---|---|
| Claude Opus 4.1 | 56% | 44% | 42% | 37% |
| GPT-5 high | 48% | 33% | 35% | 30% |
| o3 high | 45% | 28% | 29% | 25% |
| o4-mini high | 38% | 25% | 23% | 19% |
| GPT-4o | 19% | 7% | 8% | 9% |
The strong models are above parity on sub-two-hour work and fall roughly ten to twenty points across the duration axis — the time-horizon shape arrived at from the quality side, on the same benchmark that is used to argue against exponential horizon extrapolation. The 47.6% aggregate is an average over a range that starts at 56%.
Sector (Figure 10). Claude Opus 4.1 clears or approaches parity in retail trade (~56%), government (~52%) and wholesale trade (~52%) — the three the paper names — and bottoms out in information (~32%), the sector holding producers and directors, editors, journalists, audio/video technicians and film editors. The spread across sectors for a single model is roughly as wide as the spread across models within a sector.
Deliverable type (Figure 12, with each type's share of tasks):
| Deliverable | Share | GPT-4o | o4-mini high | o3 high | GPT-5 high | Claude Opus 4.1 |
|---|---|---|---|---|---|---|
| pure text | 9.9% | 5% | 15% | 21% | 23% | 14% |
| 31.3% | 9% | 18% | 25% | 30% | 45% | |
| xlsx | 23.1% | 12% | 26% | 27% | 35% | 43% |
| pptx | 5.6% | 11% | 27% | 28% | 25% | 45% |
| other | 30.1% | 15% | 34% | 43% | 46% | 48% |
This is the per-model split the vendor model-selection guidance on Cost-per-Task Over Cost-per-Token recommends, with the benchmark's own evidence behind it: Claude leads every file-producing category, GPT-5 high leads pure text. The unexpected part is that pure text is the hardest category for every model — the format most benchmarks are made of is the one where deliverables least often beat a professional's, and it is under 10% of GDPval.
Together these are jaggedness measured on paid work at three different granularities, which is the benchmark's own argument for reporting by profession rather than as a single number. Chowdhery adds a caveat against her own chart on the software-developer row: the expert baseline there was probably drawn from repository maintainers rather than generalists, so the comparison class differs from the one used for industrial or mechanical engineers. The paper's expert criteria neither confirm nor rule this out.
What the models fail at: instruction following, not knowledge#
The paper clusters expert justifications for preferring the human deliverable into three labels (Figure 8, as a share of total samples):
| Failure label | Gemini 2.5 Pro | Grok 4 | Claude Opus 4.1 | GPT-5 high |
|---|---|---|---|---|
| Instruction following | 40% | 35% | 14% | 9% |
| Formatting | 6% | 13% | 5% | 10% |
| Accuracy | 7% | 5% | 6% | 5% |
Instruction following dominates for three of four models. GPT-5 high is the exception, and the correction to the lecture's summary: its largest loss category is formatting (10%), narrowly ahead of instruction following (9%) — the lecture's "GPT-5 has fewer instruction-following errors than its peers" is right about the ranking (9% against 14/35/40%) but the page previously implied instruction following dominated everywhere. Accuracy is the smallest category for every model: the failures are not knowledge failures.
The sharpest form is a failure that reads as success — the paper's own words are that Gemini and Grok "frequently promised but failed to provide deliverables, ignored reference data, or used the wrong format", and the lecture's gloss:
the models will promise to look at the reference data but then actually not look at it… they will override it with whatever hallucination they want to come up with
With 67.7% of tasks shipping reference files, a model that says it consulted them and did not is failing the benchmark's central mechanic while producing a plausible artifact.
How bad is a failure? Experts re-graded the subset of tasks where GPT-5's deliverable lost to the human (Figure 14):
| Rating of a GPT-5 loss | Share |
|---|---|
| Model actually better (re-grader disagrees with the original) | 22.9% |
| Acceptable but subpar | 47.7% |
| Bad — unusable, not dangerous | 26.7% |
| Catastrophic — harmful or dangerously wrong | 2.7% |
The three figures as transcribed do not cleanly sum and the middle one is spoken twice with different values; carry the shape rather than the numbers. (superseded 2026-09-10: the lecture's ~50% / ~20% / ~29% were accurate to the paper's 47.7% / 22.9% / 29.4% — the hedge was unnecessary.) The denominator does need correcting, and the lecture-derived page had it wrong: these are shares of GPT-5's losses, not of all its outputs. Nearly half of what loses is still usable work; about three in ten is not; roughly one loss in forty is actively dangerous. And the 22.9% "model better" rate is itself a grader-noise measurement — the paper notes it "roughly corresponds to the level of inter-rater agreement", i.e. a fifth of recorded losses are disagreements between graders rather than model failures.
Pricing the oversight: the number that changes the cost story#
§3.2 and Appendix A.2.1 are the material the lecture could not carry and the most transferable part of the paper. The paper measures four quantities on the gold subset and then composes them into a policy:
- Human completion:
H_T= 404 minutes,H_C= $361 (validated self-reported time × median OEWS wage; the paper notes this under-estimates its experts' true market cost). - Human review of a model deliverable:
R_T= 109 minutes,R_C= $86 — read off task-monitoring telemetry, first-grade time per expert. - Model completion:
M_T,M_Cfrom empirical API latency and invoiced cost, three completions per task. - Win rate
w, which is the probability the review passes.
Under "try n times, then fix it yourself" the expected cost is (M_C + R_C)·(1−(1−w)ⁿ)/w + (1−w)ⁿ·H_C, and the results (Table 2, verified against the PDF at compile time) are the point:
| Model | Win rate | Speed: naive | try 1× | try n× | Cost: naive | try 1× | try n× |
|---|---|---|---|---|---|---|---|
| gpt-4o | 12.5% | 327× | 0.87× | 0.46× | 5172× | 0.90× | 0.53× |
| o4-mini | 29.1% | 186× | 1.02× | 1.06× | 1265× | 1.06× | 1.22× |
| o3 | 35.2% | 161× | 1.08× | 1.28× | 480× | 1.13× | 1.47× |
| gpt-5 | 39.0% | 90× | 1.12× | 1.39× | 474× | 1.18× | 1.63× |
Three readings:
- The naive ratio is a fiction, and the paper says so. A model that is 474× cheaper than the expert in raw API spend is 1.63× cheaper once you pay an expert $86 to check the output and pay the full $361 again whenever it fails. Review is 27% of the task's own duration and it is charged on every attempt, so a three-order-of-magnitude advantage is spent almost entirely on oversight. This is Verification as the New Bottleneck with a receipt.
- Below a win-rate threshold the loop is worse than doing it yourself. GPT-4o at 12.5% lands at 0.46× speed and 0.53× cost — resampling a model that usually fails burns review time and then costs the human the task anyway. The sign of the benefit flips on
w, which makes win rate a deployment parameter, not just a leaderboard number. - The lecture's cost claim was correct and this page was wrong to discard it.
A cost/speed claim that should not be carried at all… those two are not reconcilable as transcribed; the ASR is damaged here.(superseded 2026-09-10.) Chowdhery's "about 1.6× in cost improvement and about 1.4 in speed improvement" is GPT-5's try-n-times row exactly (1.63× / 1.39×), and her "less than 10% of the human expert salary" is the naive column of the same table read loosely — the raw sampling cost is far below 10% of the wage bill (474× is ~0.2%). They were two different rows, not a contradiction.
What the accounting excludes, by the paper's own admission: the time to review a human deliverable (which real workplaces also pay), the possibility that the human deliverable is bad, and the cost of catastrophic mistakes — which the failure analysis above prices at 2.7% of losses, and which are exactly the errors whose cost is not proportional to the task. Footnote 9 argues the model is over-penalized (win rate should rise as the professional adapts the prompt, review time should fall with familiarity). And the scope limit that is a COI issue: cost estimates were obtained for OpenAI models only — "we were not able to obtain cost estimates for Claude, Gemini, and Grok" — so the model that wins the headline has no row in the savings table. Compare Expenditure Horizon, METR's dollar-denominated metric, which puts the same human-versus-agent comparison on a money axis for a single optimization problem.
What elicitation buys: reasoning effort, prompting, scaffolding#
The paper runs three non-interactive levers (§3.4, Figure 9), all on the same gold subset and the same graders:
| Lever | Result |
|---|---|
| Reasoning effort, o3 | low 29.8% → medium 30.8% → high 34.1% |
| Reasoning effort, GPT-5 | low 32.7% → medium 35.6% → high 38.8% (+6.1 pp) |
| Prompt tuning, GPT-5 high | 38.8% → 43.1% (+4.3 pp; the prose rounds this to "5 percentage points") |
| Scaffolding | GET requests enabled in the container; best-of-N sampling at N=4 with a GPT-5 judge |
The prompt is a formatting-and-self-checking discipline, not task knowledge: render every visual deliverable to PNG and look at it, run programmatic checks for blank pages and overflow, use LibreOffice for PDFs, avoid non-standard characters, cap deliverable text at four sentences. Its measured effects are the striking part — black-square artifacts, previously affecting over half of generated PDFs, went to zero; egregious PowerPoint formatting errors fell 86% → 64%; and agents inspecting their own deliverables with their multimodal capability went 15% → 97%. Almost all of the gain comes from making the agent look at what it produced.
This is the paper's own evidence that a meaningful slice of the gap to the expert is protocol rather than capability, and it is corroborated from outside: HarnessBank's evolved harness moves a GDPval domain 43.7 → 52.9 Pass@1 (+9.2, credited on a sealed held-out split) — see Agent-Authored Harness Optimization. What remains untested is the specific protocol the benchmark excludes by construction: a human correcting the first attempt.
Grading: blinded experts, and a grader validated against the human ceiling#
Grading the 220 gold tasks is blinded expert pairwise comparison — occupational experts see the request and reference files and rank two or more unlabeled deliverables, with written justifications. Each comparison takes over an hour on average, which is why the automated grader exists.
The automated grader is GPT-5-high, released publicly at evals.openai.com. Its validation is reported the right way — against the human ceiling rather than against 100%:
- Human–automated agreement: 65.7%
- Human inter-rater agreement: 70.8%
Agreement is defined as E[1 − |H − A|] over ordinal scores in {0, 0.5, 1}; the paper explains it declines Cohen's/Fleiss's κ and Krippendorff's α because the scores are three-point ordinals. That framing — report the grader relative to what humans achieve against each other, not against a fictional perfect label — is the same move Usage-Telemetry Classifier Validation argues for on occupation classifiers, executed by a different lab in a different setting.
Two limitations the paper volunteers, and both matter for LLM-Judge Validation:
- Self-preference, measured in the field. The GPT-5-high grader "shows lower correlation with human expert graders when assessing outputs from capable OpenAI models", cited to Panickssery et al. (2024) on models favoring their own generations. This is a vendor reporting that its own judge is least trustworthy on its own models — and both agreement metrics are highest for weak models, whose outputs are easy to tell apart from a human's.
- 12 of 220 tasks are ungradable by the automated grader: tasks needing internet access, non-Python execution (three software-developer tasks), specific font packages, or speech-to-text. The public automated service is therefore a proxy on a subset, and the paper says human grading remains the recommended method.
The finding that generalizes: the benchmark measures a low-context expert#
GDPval's tasks are, by construction, one-shot and precisely specified — all the context a professional carries in their head is written into the prompt, and there is no back-and-forth in which a human corrects the first attempt. The paper lists this as a limitation in its own §5. That design choice is what makes the benchmark scorable, and it is also what it cannot measure.
The ablation is Appendix A.2.7: rewrite the prompts to omit where the data lives, how to approach the problem, and what the deliverable should look like — down to 42% of the original token length — and re-grade GPT-5-high with human experts. The result (Figure 15):
| Prompt | Wins + ties | Wins only |
|---|---|---|
| Original | 47.7% | 43.3% |
| Concise / underspecified | 44.3% | 39.8% |
A 3.4-point drop, and the paper flags that this arm ran on an earlier version of the gold subset, so its 47.7% baseline is not the 38.8% in the main text and the two must not be differenced. The metric moves little; the qualitative failure is larger than the metric suggests — "the models struggled to figure out context", or in the lecture's phrasing, "the models actually struggle to figure out what to work on." Chowdhery's reading:
real work is often context heavy… humans are basically architecting what is the set of problems and then the model can go solve it
This is Context Advantage, Not Taste measured on paid work rather than argued: the residual human contribution is an information asymmetry, and GDPval prices the half of the job that survives when the asymmetry is deliberately erased. It lands on the same side as Returns to Expertise in Agentic Coding, and the lecture connects it to METR's contractor-versus-maintainer finding directly — the model behaves like a smart newcomer with no context rather than like the embedded expert.
Limitations, and three inconsistencies inside the paper#
The paper's own §5: 44 occupations and ~30 tasks each is "a limited, initial cut"; only self-contained digital knowledge work (no manual labor, no tacit knowledge, no PII, no proprietary tools, no communication between individuals); one-shot and non-interactive; grader limitations; and the expense of running it at all.
Three internal inconsistencies found at compile time, none of them parse damage (all seven tables were re-verified against pdftotext -layout):
- Reference-file counts. §1 claims "up to 17 reference files in the gold subset, and 38 in the full set"; Table 5, captioned gold set, gives a maximum of 38. Either the caption or the intro is wrong. The safe figure is the mean (1.92) and the 67.7%-of-tasks share.
- Win rates differ between the headline figure and the cost table. Figure 5 gives GPT-4o 12.4 / o4-mini 27.9 / o3 34.1 / GPT-5 38.8; Table 2's
wcolumn gives 12.5 / 29.1 / 35.2 / 39.0. The paper never reconciles them, so the savings ratios rest on a slightly different aggregation from the headline. - Three different "how long does a task take" numbers: 7 hours (intro), 9.49-hour gold-subset mean (Table 3), and
H_T= 404 minutes ≈ 6.7 hours (the cost model). They are different estimands — a rounded headline, a task-characteristics summary, and the validated input to the savings formula — but nothing in the paper says so.
Where GDPval sits in the capability space (2026-09-22)#
The first external, quantitative answer to "does GDPval measure something the other benchmarks do not" arrives from a psychometric audit rather than from OpenAI: Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author unrefereed preprint) treats twelve benchmarks as items and 421 model configurations as respondents on a hash-pinned Artificial Analysis snapshot, with every hypothesis carried on the 96-model grid scored on all twelve. Four results land on this page.
GDPval is agentic, not sui generis. Under a three-factor over-extraction of the date-adjusted scores, GDPval loads 0.84 on an agentic/work-realistic factor alongside Terminal-Bench v2.1 (1.01), Terminal-Bench Hard (0.82), AA-Omniscience (0.69), τ³-Banking (0.54) and τ²-Bench (0.50) — not on the academic-knowledge factor (GDPval's loading there is 0.13) and not on the physics/hard-reasoning one (−0.04). Whatever separates GDPval from GPQA is the same thing that separates a terminal agent from a multiple-choice test, and it is shared with two benchmarks the paper's own taxonomy calls non-economic.
Under the pre-specified rule there is no economic factor at all. Horn's parallel analysis retains exactly one factor on the date-adjusted data, so the separation above is an exploratory over-extraction that no dimensionality criterion selects; the author reports the structural hypothesis as not supported as specified. GDPval's bootstrap loading interval is [0.44, 0.99] — it excludes zero and tells you nothing finer.
But GDPval carries information a general index does not. Held out entirely and predicted from the other eleven benchmarks with factors re-estimated inside every fold, a three-factor representation beats a single mean-score index by ΔMSE +0.026, 95% CI [+0.010, +0.044], and the best model reaches out-of-fold R² = 0.88. So a buyer who reads only an intelligence index is losing real, if incremental, information about professional-task quality — which is the quantitative case for the separate column this benchmark exists to justify.
And the scale story inverts. SHAP attribution on the matched gradient-boosting learner ranks the agentic factor first, then a reasoning flag, then the hard-reasoning factor, with log-params and open-weights status contributing almost nothing. Prior work correlates a general capability factor with parameter count at ~0.6; here the scale signal reaches GDPval through the factors rather than alongside them. The largest single residual is interpretable rather than diagnostic — Grok 4.3 Non-reasoning beats its economic prediction by +1.14 z-units, which fits a benchmark rewarding agentic behaviour the others capture only indirectly.
The caveat that bites this page specifically. The GDPval column in that snapshot is the Artificial Analysis GDPval Elo board, not the paper's own win rate, and Artificial Analysis records that board as having been silently rescored between a v1 and a v2 that are not comparable. Zhu names no version. Nothing in the covariance analysis depends on it, but no number above can be lined up against a win rate on this page.
Connections#
- Task Time-Horizon Scaling — the metric this one is designed to argue with: duration-at-fixed-reliability rising exponentially versus quality-against-an-expert rising roughly linearly, on overlapping work. The paper's linearity claim is a three-point OpenAI-only fit ending at 38.8%, and its own duration gradient (56% at 0–2 h down to 37% at 8+ h for the best model) is the horizon shape re-measured on the quality axis
- Economic Benchmark Construct Validity — the external construct-validity audit of the column GDPval anchors, and the first quantitative account of what this benchmark measures that does not come from its authors. GDPval loads 0.84 on an agentic factor shared with Terminal-Bench, forms no distinct economic factor under the pre-specified rule, and still adds ΔMSE +0.026 [+0.010, +0.044] over a single intelligence index out of sample — incremental validity without a separate construct. It also supplies the reason a small GDPval gap between models released months apart should be date-adjusted before it is read
- Expenditure Horizon — the nearest sibling and the closest thing to a joint method: METR prices agent-versus-human on a dollar axis for one optimization problem, GDPval prices it across 44 occupations by composing human wage cost, measured expert review time, and API cost into a savings ratio. GDPval's version is the one that includes oversight explicitly, and finds it eats a 474× advantage down to 1.63×
- Verification as the New Bottleneck — the same bottleneck with a price tag: 109 minutes of expert review against a 404-minute task, charged on every resample, is what turns the naive cost advantage into a rounding error
- DRACO Benchmark — the methodological twin and the COI mirror. Both are vendor-built benchmarks of real professional work graded by expert judgment; on DRACO the vendor's own product wins every domain, on GDPval the vendor's own model loses to a competitor by 8.8 points. That asymmetry is the cheapest available evidence about which vendor benchmark to discount
- Production-Sourced Evaluation — the adjacent-but-distinct sourcing method. DRACO mines de-identified production traffic; GDPval commissions professionals to submit their own real work product, then puts each task through five expert reviews. Neither synthetic nor hand-authored-from-imagination nor production-sampled — a third route that buys occupational representativeness and a wage-denominated value per task, at the cost of a distribution chosen by the benchmark's authors rather than observed
- LLM-Judge Validation — GDPval's grader validation is a field instance of two things that page tracks: agreement reported against the human inter-rater ceiling (65.7% vs 70.8%) rather than against 100%, and a measured self-preference effect — the GPT-5-high grader agrees less with humans precisely on capable OpenAI models' outputs
- Usage-Telemetry Classifier Validation — the same ceiling-relative reporting discipline in a different measurement problem; two labs independently arriving at "score the automated rater against what humans achieve against each other"
- Agent-Authored Harness Optimization — GDPval as a domain in HarnessBank's sealed-split harness evolution (43.7 → 52.9 Pass@1, +9.2), which together with this paper's own prompt-tuning and best-of-N results makes a substantial slice of the gap to the expert protocol rather than capability
- Deep Research Agents — the third benchmark in the same lecture (DeepScholar-Bench) is the follow-up to this page's reference-file failure mode: GDPval shows models skipping supplied references; DeepScholar-Bench measures what happens when they must find the references themselves
- Context Advantage, Not Taste — the underspecification ablation is this thesis as an experiment, now with numbers: prompts cut to 42% of their length move GPT-5's win-or-tie rate 47.7% → 44.3%, while the qualitative failure ("struggled to figure out context") is larger than 3.4 points
- Returns to Expertise in Agentic Coding — the same asymmetry from the human side; GDPval's baseline is a professional averaging 14 years in the occupation whose advantage is largely context, and the benchmark is built to hand that context over
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — a fifth instrument for the same family of questions, measuring something the other four do not: not what an LLM could do, what workers say it does, or what telemetry observes, but whether a deliverable actually beats a professional's. Its occupation and sector selection is itself an exposure-measurement design choice — GDP contribution first, then a ≥60% digital-task share
- Task Saturation: Broad but Shallow AI Diffusion — the complementary cut of the same labor market: ATLAS asks what share of an occupation's tasks AI touches at all, GDPval asks how good the output is on the highest-wage digital slice of nine sectors, selected by wage mass rather than by observed usage
- Cost-per-Task Over Cost-per-Token — vendor model-selection guidance cites the GDPval-AA derivative for knowledge work; the per-deliverable-type split here (Claude leading pdf/xlsx/pptx, GPT-5 leading pure text) is that advice with the benchmark's own numbers behind it
- The Open-Weight Frontier Gap — the GDPval-AA Elo board is where the open/closed agentic gap is measured at 61 Elo; this is what the underlying benchmark actually asks
- Measuring Beyond Accuracy Saturation — the same rejection of headline accuracy from the other direction: that page re-instruments a saturated benchmark along non-accuracy axes, this one changes the object to economically valuable work graded against a human. GDPval's own no-upper-limit argument is that a win rate never saturates because the baseline can be replaced
- Failures That Look Like Success — GDPval's dominant failure mode is this class: instruction-following failures at 35–40% of samples for Grok and Gemini, models that promise deliverables and do not provide them, and reference data ignored while the artifact still looks right
- Jagged Intelligence (Ghosts, Not Animals) — jaggedness measured at three granularities on paid work: occupation, sector (32% to 56% for one model), and deliverable file type
- OpenAI — the benchmark's author, and the conflict of interest
- Aakanksha Chowdhery — the lecturer whose CS329A session was this page's only source until 2026-09-10, and whose slide-read figures the paper largely vindicates
- CS329A: Self-Improving AI Agents (Stanford) — the course; lecture 8 pairs this benchmark with METR's time horizons and DeepScholar-Bench as three axes of agentic evaluation
- Cognitive Capability Profiling for Task Suitability — the capability-decomposed counterpart, and the comparison neither paper makes. GDPval asks whether a model's deliverable beats a 14-year professional's on real work product, scored as a blinded pairwise win rate, and has no capability decomposition at all; Prunty et al. infer a model's cognitive profile across eight dimensions from 19,535 rubric-annotated benchmark items, map it against what 410 workers say their activities require, and never observe an output. One measures output against the incumbent; the other decomposes capability and validates against nothing. Running their six profiled systems on this benchmark's 220-task open gold subset and correlating win rate against per-activity suitability is the obvious unbuilt validation for both pages — and both take their occupational vocabulary from O*NET, so the crosswalk already exists
Open Questions#
- GDPval scores one-shot delivery with no iteration. How much of the gap to the expert closes when the model is allowed the back-and-forth a real assignment gets — and is that gap capability or protocol? Partially answered 2026-09-10 by GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, on the non-interactive half. Three levers the paper does test move it materially without any human in the loop: reasoning effort (GPT-5 low → high, 32.7% → 38.8%), a generic formatting-and-self-check prompt (38.8% → 43.1%, with self-inspection of deliverables going 15% → 97% and black-square PDF artifacts eliminated), and scaffolding (best-of-N=4 with a judge, container GET access). HarnessBank's sealed-split harness evolution adds +9.2 points on a GDPval domain from outside the paper. So a substantial slice is protocol. The interactive half remains untested — the paper lists one-shot non-interactivity as its own limitation and promises future versions — and the falsifiable form sharpens: run the gold subset with a bounded number of expert correction turns and compare the win-rate lift against the 4.3-point prompt-tuning lift, which is the non-interactive ceiling on the same graders.
- The headline comparison is system-level, not model-level: Claude Opus 4.1 was sampled through its consumer UI to enable file creation, while OpenAI models ran via API with web search, a code interpreter and preinstalled libraries. How much of the 8.8-point gap between Claude Opus 4.1 and GPT-5 high is the model, and how much is the sampling surface? Falsifiable directly — run both through matched API scaffolds on the open 220-task gold subset.
- The savings analysis covers OpenAI models only — the paper could not obtain cost estimates for Claude, Gemini or Grok — so the model with the highest win rate has no row in the table where win rate is the parameter the payoff turns on. Does the try-n-times ratio for a 47.6% model clear GPT-5's 1.63×, and at what price per deliverable?
Resolved Questions#
The paper itself is not in this corpus — every figure here is a lecturer's slide reading. What do GDPval's published numbers actually say, and how much of the ASR-damaged quality distribution and cost claim survives contact with the source?Answered 2026-09-10 by GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (arXiv 2510.04374 v1,empirical). Nearly all of it survives. The quality distribution was accurate as spoken (47.7% acceptable-but-subpar / 22.9% model-better / 29.4% bad-or-catastrophic) but its denominator was wrong on this page — those are shares of GPT-5's losses, not of all outputs. The cost claim was also accurate and was wrongly discarded here as irreconcilable: "1.6× cost, 1.4× speed" is GPT-5's try-n-times row (1.63× / 1.39×) and "under 10% of the expert's salary" is the naive column of the same table. Three things did not survive: the sector rule (>5% each, not a top-5% slice), the O*NET digital filter (a 60% per-occupation inclusion threshold, not a share of tasks kept), and the expert bar (4-year floor, 14-year mean, not "10+ years"). The largest correction is not from the lecture at all but from this page's own reading: the roughly-linear trend is a three-point OpenAI-only fit ending at 38.8%, and 47.6% is a competitor's score on a different chart with no time axis.
Sources#
-
One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation, arXiv 2608.29420, 2026-08-29, 25pp,
empirical. Cited here for the section above: §5.3 and Appendix E Table 3 (the date-adjusted three-factor loadings, GDPval 0.84 / 0.13 / −0.04, and the one-factor retention that makes the split exploratory), §5.4 and Table 1 (the ΔMSE of +0.026 [+0.010, +0.044] and the ladder it sits in), Appendix M (out-of-fold R² = 0.88, the SHAP ordering, the Grok 4.3 Non-reasoning +1.14 residual), and Appendix F (GDPval scored on 117 of 548 raw configurations, 112 retained — the sparsest-but-one column in the battery). Tables 1, 3 and 4 were reconciled againstpdftotext -layoutand are exact. Scope: the paper fixes itself to internal validity and says nothing about whether GDPval predicts any real economic outcome. Full treatment on Economic Benchmark Construct Validity -
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan, Dias, Proehl, Kim, Wang, Watkins, Posada Fishman, Aljubeh, Thacker et al. (19 authors, all OpenAI), GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, arXiv 2510.04374 v1, 2025-10-05, 29pp (
empirical). The primary source for everything on this page except where attributed to the lecture: §2 (sector and occupation selection, expert recruitment, task creation, the review pipeline), §2.5 and A.6 (grading protocol, automated-grader agreement and its self-preference and ungradable-task limitations), §3.1 and Figures 5–6 (headline win rates; the OpenAI-only time series), §3.2 and A.2.1 with Table 2 (the speed/cost model and its ratios), §3.3 and Figure 8 (failure clustering), §3.4 and Figure 9 (reasoning effort, prompt tuning, scaffolding), §5 (limitations), A.2.2–A.2.6 (sector, occupation, deliverable, duration and failure-severity breakdowns), A.2.7 and Figure 15 (the underspecification ablation), A.4 Tables 3–7 (task characteristics), A.7 (the O*NET digital-task methodology and its Acemoglu–Autor validation). Baseline-lane ingest, deliberately outside the usual freshness window, taken because this page had been written entirely from a lecture. An ICLR 2026 camera-ready exists at proceedings.iclr.cc and was not ingested; nothing here should be quoted as the camera-ready's text. Parse notes: all seven tables re-verified cell-by-cell againstpdftotext -layoutat compile time (Table 2 p.13, Tables 3–6 p.19, Table 7 p.20, Table 1 p.29) — exact match, no collapse, shift or weld. Table 7's run-together header (%, gold set%, full set) is faithful to the PDF's own rendering; the first data column is the gold set. Figure 5 is image-only and was transcribed at ingest; that transcription plus Figures 4b, 6–15 were viewed directly at compile time before any figure-derived number below was quoted. COI: OpenAI authored the benchmark, selected the occupations, paid the experts, defined the metric, and supplied both four of the seven evaluated models and the GPT-5-high automated grader — but a competitor's model wins the headline, and the paper reports its own grader's self-preference against itself. Where the COI does bite is the cost analysis, which covers OpenAI models only -
CS329A Self-Improving AI Agents — Part 8: Agentic Evaluations and Long-Horizon Tasks — Aakanksha Chowdhery, Stanford CS329A lecture 8, delivered 2025-11-17, published 2026-08-03 (
practitioner-opinion, YouTube auto-caption transcript, ~12.9k words). This page's original and only source until 2026-09-10, now demoted to a secondary reading and retained for what the paper does not contain: the METR-versus-GDPval framing quoted above, the example-task list, the per-occupation near-parity commentary and the repository-maintainer caveat on the software-developer baseline, and the interpretive gloss on the underspecification result. Figures were read off slides through ASR and are superseded above wherever the paper differs; the strike-throughs on this page record which lecture-derived claims held and which did not. No conflict of interest — the benchmark is OpenAI's and neither CS329A instructor is an author
Cited by 26
- CS329A: Self-Improving AI Agents (Stanford)×4
GDPval (OpenAI) · win rate against a practising professional on real paid work — Gdpval Benchmark ·…
- Artificial Analysis×3
GDPval-AA (Elo, v1 and v2) · Gdpval Benchmark, Claude Opus 4 8, Kimi, Open Weight Frontier Gap · an…
- Task Time-Horizon Scaling×3
Gdpval Benchmark — the deflationary counterweight taught alongside this metric: win rate against a…
- Aakanksha Chowdhery×2
Gdpval Benchmark — lecture 8's middle third, and until 2026-09-10 the wiki's only account of the…
- Agent-Authored Harness Optimization×2
Gdpval Benchmark — the domain in HarnessBank's suite whose primary benchmark this wiki now holds,…
- Cognitive Capability Profiling for Task Suitability×2
Gdpval Benchmark — the outcome-measured counterpart, and the comparison this paper never makes.…
- Context Advantage, Not Taste×2
Gdpval Benchmark — the asymmetry run as an experiment on paid work, and since 2026-09-10 with the…
- Cost-per-Task Over Cost-per-Token×2
Gdpval Benchmark — what the GDPval-AA row in the model-selection table above is built on, and where…
- DRACO Benchmark×2
gdpval real world economically valuable tasks — Patwardhan et al. (19 authors, OpenAI), arXiv…
- Expenditure Horizon×2
Gdpval Benchmark — the nearest sibling on the dollar axis, and the one that prices the piece this…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated×2
Gdpval Benchmark — an instrument outside this page's four, and outside Steele & Cruz's seven,…
- Failures That Look Like Success×2
Gdpval Benchmark — the class as the dominant failure mode of an economically-framed benchmark…
- Jagged Intelligence (Ghosts, Not Animals)×2
Gdpval Benchmark — jaggedness measured on paid work at three granularities, and the reason its…
- LLM-Judge Validation×2
gdpval real world economically valuable tasks — Patwardhan et al. (19 authors, OpenAI), arXiv…
- Open Questions Backlog×2
Gdpval Benchmark ×2 (oldest 19d) — The headline comparison is system-level, not model-level: Claude…
- Production-Sourced Evaluation×2
Gdpval Benchmark — the third sourcing route, adjacent to this page's and distinct from both of its…
- Returns to Expertise in Agentic Coding×2
Gdpval Benchmark — the benchmark that deliberately erases the premium, and since 2026-09-10 with…
- Task Saturation: Broad but Shallow AI Diffusion×2
gdpval real world economically valuable tasks — Patwardhan et al. (19 authors, OpenAI), arXiv…
- Usage-Telemetry Classifier Validation×2
Gdpval Benchmark — the same reporting discipline reached independently by a different lab on a…
- Claude Opus 4.8
The GDPval-AA row is the v1 board as reported in 4.8's own system card. Artificial Analysis later…
- Deep Research Agents
Read against GDPval from the same lecture, the two failures are one failure at different distances:…
- Economic Benchmark Construct Validity
Gdpval Benchmark — the economic benchmark this snapshot scores highest and predicts best, sitting…
- Measuring Beyond Accuracy Saturation
Gdpval Benchmark — the third response to "traditional benchmarks are saturating", after…
- Evals & Benchmarks
Gdpval Benchmark — OpenAI's benchmark of real, economically valuable knowledge work: 1,320 tasks…
- The Open-Weight Frontier Gap
Gdpval Benchmark — the benchmark under the GDPval-AA Elo board this page differences: real…
- OpenAI
Gdpval Benchmark — OpenAI's benchmark for whether a model's deliverable beats a practising…
Related articles
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Production-Sourced Evaluation
Building benchmarks from de-identified real production usage rather than synthetic or hand-authored tasks; DRACO's cent…
- LLM-as-a-Judge
Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…
