Sources#
Summary#
GDPval is OpenAI's benchmark for whether a model can do real work that people are paid for. Instead of scoring an answer against a key, it puts the model's deliverable next to a deliverable produced by an industry professional with more than a decade of experience in that occupation and asks graders which is better — a pairwise win rate, reported as wins-only and as wins-plus-ties.
The wiki has been quoting GDPval's descendants for months without holding the original: GDPval-AA (the Artificial Analysis Elo board built on it) is cited on Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Kimi (Moonshot AI), The Open-Weight Frontier Gap and Cost-per-Task Over Cost-per-Token, among others. This page is the ancestor those numbers descend from.
The derivative board has been rescored, and the versions are not comparable. Artificial Analysis's original GDPval-AA board (quoted in Opus 4.8's system card) and its v2 successor (quoted in Opus 5's, and in Kimi K3's and The Open-Weight Frontier Gap's 61-Elo gap) put the same model — Opus 4.8 — at 1890 and 1593 respectively. v2 rebuilt the run: 220 tasks from the gold set, shell access and web browsing in an agentic loop, Elo from blind pairwise comparisons. So a GDPval-AA number is only meaningful with its board version attached, and v1 and v2 Elo must never be differenced against each other — the same caution The Open-Weight Frontier Gap applies when it declines to difference an IRT cyber Elo against an Arena gap.
Provenance and confidence. Everything here comes from one teaching account — Aakanksha Chowdhery's lecture 8 of CS329A (delivered 2025-11-17, practitioner-opinion), read off slides through a YouTube auto-caption transcript. The benchmark is OpenAI's, published around September 2025; neither instructor is an author. Treat every figure below as a lecturer's reading of a slide, not as a quotation of the paper — several are visibly ASR-damaged and are flagged where they are.
The design#
| As described in lecture | |
|---|---|
| Sectors | 9, selected as the top ~5% of GDP contribution — real estate, government, professional/scientific/technical services, healthcare, finance, retail & wholesale, information (incl. film and video editing), manufacturing |
| Occupations | 44 |
| Tasks | ~1,320 (ASR renders it "30 20"; reconstructed), of which ~220 form an open gold set released publicly |
| Task source | Real work supplied by practising professionals with 10+ years' experience in the occupation |
| Task frame | O*NET occupation-task taxonomy, filtered to the ~60% of tasks that are computer-based / digital |
| Duration | ~7 hours average expert completion time; some tasks run to weeks |
| Modality | Text and multimodal — CAD, video, audio, spreadsheets, presentations |
| Value | ~$400 average per task in the gold subset, with a long tail of higher-value work |
| Reference files | ~70% of tasks require interacting with supplied reference files |
| Specification quality | ~89% vetted by experts as well specified — i.e. the benchmark deliberately removes ambiguity as a failure cause |
Representative tasks named in lecture: design a 3D model of a cable-reel stand for an assembly line (manufacturing engineer); build a competitive landscape (financial analyst); assess images and draft a consultation note (registered nurse); cut an intro reel from a script (film/video editor); draft a response to a dissatisfied customer (customer service); plan an itinerary for a family of four (concierge); find pricing inconsistencies across purchase orders (audit); design a property sales brochure (real estate agent); optimize a vendor-fair layout (recreation worker).
The headline trend: linear, where METR's is exponential#
| Model | ~Date | Win rate vs. expert |
|---|---|---|
| GPT-4o | 2024 | ~12.4% |
| (intermediate models) | 2024–25 | ~25%, then low-to-mid 30s |
| Claude Opus 4.1 | 2025 | ~47.6% |
Chowdhery states the comparison explicitly and it is the most consequential thing this benchmark does to the rest of the wiki: "that's a very different trend compared to METR where we were talking about this doubling trend every seven months. It's more of a linear trend roughly compared to the exponential trend that METR was talking about."
The two are not the same measurement — one is task duration at fixed reliability, the other is output quality against a human at fixed task — and the lecture says so when a student presses. But the deflationary reading is deliberate: the exponential curve invites the extrapolation one hour now, a couple of days next, a couple of weeks after that, and GDPval's answer is that on work that actually takes several hours to weeks, the models are only so useful, and the usefulness is broken down by profession rather than aggregated into a single horizon. See Task Time-Horizon Scaling.
Two structural qualifiers the lecture attaches to the win rate:
- Duration gradient. Models do better on shorter tasks (up to a few hours) and decline as tasks lengthen — the same shape as the time-horizon curve, arrived at from the quality side.
- Occupation gradient. Near-parity with human experts appears in a specific and unglamorous list: counter and rental clerks, real estate brokers, shipping/receiving/inventory clerks, buyers and purchasing agents, computer and IT systems managers, software developers; then administrative services managers, compliance officers, medical and health services managers, personal financial advisors, customer service representatives; then first-line retail supervisors, non-retail sales workers, news analysts, and wholesale/manufacturing sales representatives. The aggregate number is an average over a distribution this wide, which is jaggedness measured at occupation granularity.
Chowdhery adds a caveat against her own chart on the software-developer row: the expert baseline there was probably drawn from repository maintainers rather than decade-of-experience generalists, so the comparison class differs from the one used for, say, industrial or mechanical engineers.
What the models fail at: instruction following, not knowledge#
The failure analysis is the part of GDPval that most resembles the rest of this wiki's evidence. The dominant failure category across models is instruction following, not competence — and its sharpest form is a failure that reads as success:
the models will promise to look at the reference data but then actually not look at it… they will override it with whatever hallucination they want to come up with
Formatting errors are the second cluster. GPT-5 is reported to have fewer instruction-following errors than its peers. Given that ~70% of tasks ship reference files, a model that says it consulted them and did not is failing the benchmark's central mechanic while producing a plausible artifact.
Model-specific strengths (late 2025, as observed by the benchmark): Claude tends to win on aesthetics, document formatting, and understanding PDFs, spreadsheets and presentations; GPT-5 tends to win on instruction following, correct calculations, and text-only tasks. The practical instruction the lecture draws — choose the model per task type — is the same selection discipline the vendor guidance arrives at from the cost side.
A quality distribution, hedged. For GPT-5 the lecture reports roughly half of outputs as acceptable but subpar, ~20% where the model's output would genuinely be better than the expert's, and ~29% bad or catastrophic — with acknowledged disagreement between human graders. The three figures as transcribed do not cleanly sum and the middle one is spoken twice with different values; carry the shape (a large mediocre middle, a real superhuman tail, a substantial unusable tail) rather than the numbers.
A cost/speed claim that should not be carried at all. The lecture states a best-of-n self-repair loop yields "about 1.6x in cost improvement… and about 1.4 in speed improvement relative to an unaided expert", then immediately that successful model work costs "less than 10% of the human expert salary". Those two are not reconcilable as transcribed; the ASR is damaged here. What survives is the direction — that iterated sampling with self-repair improves both axes, and that on the subset where the model succeeds it is much cheaper than the expert — with no usable magnitude.
The finding that generalizes: the benchmark measures a low-context expert#
GDPval's tasks are, by construction, one-shot and fully specified: all the context a professional carries in their head is written into the prompt, and there is no iterative back-and-forth in which a human corrects the first attempt. That design choice is what makes the benchmark scorable, and it is also what it cannot measure.
The lecture's ablation is the evidence: underspecify the prompts — strip context back out — and the win rate falls a few points, but the qualitative failure is larger than the metric suggests. "The models actually struggle to figure out what to work on." Chowdhery's reading:
real work is often context heavy… humans are basically architecting what is the set of problems and then the model can go solve it
This is Context Advantage, Not Taste measured on paid work rather than argued: the residual human contribution is an information asymmetry, and GDPval prices the half of the job that survives when the asymmetry is deliberately erased. It also lands on the same side as Returns to Expertise in Agentic Coding — and the lecture connects it to METR's contractor-versus-maintainer finding directly, arguing the model behaves like a smart newcomer with no context rather than like the embedded expert.
Connections#
- Task Time-Horizon Scaling — the metric this one is designed to argue with: duration-at-fixed-reliability rising exponentially versus quality-against-an-expert rising roughly linearly, on overlapping work. The lecture treats them as complementary axes and uses this one to block the naive extrapolation from the other
- Deep Research Agents — the third benchmark in the same lecture (DeepScholar-Bench) is the follow-up to this page's reference-file failure mode: GDPval shows models skipping supplied references; DeepScholar-Bench measures what happens when they must find the references themselves, and no system exceeds ~19%
- Context Advantage, Not Taste — the underspecification ablation is this thesis as an experiment: erase the human's information advantage from the prompt and the model stops knowing what to work on, while staying capable of the execution
- Returns to Expertise in Agentic Coding — the same asymmetry from the human side; GDPval's expert baseline is a 10-year professional whose advantage is largely context, and the benchmark is built to hand that context over
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — a fifth instrument for the same family of questions, and one that measures a different thing than the other four: not what an LLM could do, what workers say it does, or what usage telemetry observes, but whether a deliverable actually beats a professional's. Occupation-level exposure rankings and this win-rate ranking are checkable against each other
- Cost-per-Task Over Cost-per-Token — vendor model-selection guidance cites the GDPval-AA derivative of this benchmark for knowledge work; the per-model strength split here (Claude on formatting and documents, GPT-5 on instruction following and calculation) is the same advice with the benchmark's own evidence behind it
- The Open-Weight Frontier Gap — the GDPval-AA Elo board is where the open/closed agentic gap is measured at 61 Elo; this is what the underlying benchmark actually asks
- Measuring Beyond Accuracy Saturation — the same rejection of headline accuracy from the other direction: that page re-instruments a saturated benchmark along non-accuracy axes, this one changes the object to economically valuable work graded against a human. Both treat "accuracy on a task basket" as the thing to move past
- Failures That Look Like Success — GDPval's dominant failure mode is this class: a model that states it consulted the reference files, did not, and returns a well-formatted plausible artifact
- OpenAI — the benchmark's author
- Aakanksha Chowdhery — the lecturer whose CS329A session is this page's only source
- CS329A: Self-Improving AI Agents (Stanford) — the course; lecture 8 pairs this benchmark with METR's time horizons and DeepScholar-Bench as three axes of agentic evaluation
Open Questions#
- The paper itself is not in this corpus — every figure here is a lecturer's slide reading. What do GDPval's published numbers actually say, and how much of the ASR-damaged quality distribution and cost claim survives contact with the source?
- GDPval scores one-shot delivery with no iteration. How much of the ~48% gap to the expert closes when the model is allowed the back-and-forth a real assignment gets — and is that gap capability or protocol?
Sources#
- CS329A Self-Improving AI Agents — Part 8: Agentic Evaluations and Long-Horizon Tasks — Aakanksha Chowdhery, Stanford CS329A lecture 8, delivered 2025-11-17, published 2026-08-03 (
practitioner-opinion, YouTube auto-caption transcript, ~12.9k words). The GDPval section: sector/occupation/task counts, the O*NET digital-task filter, task characteristics and example tasks, the win-rate trend from GPT-4o to Claude Opus 4.1, per-model strengths, the instruction-following failure analysis, the GPT-5 quality distribution, the underspecification ablation, and the per-occupation near-parity lists. Figures are read off slides through ASR: the task count, the quality distribution and the cost/speed claim are each flagged in the text above as damaged or reconstructed. No conflict of interest — the benchmark is OpenAI's and neither CS329A instructor is an author
Cited by 16
- CS329A: Self-Improving AI Agents (Stanford)×3
Gdpval Benchmark — lecture 8's middle third, and the wiki's anchor for a benchmark it had only been…
- Aakanksha Chowdhery×2
Her fourth and final solo lecture (delivered 2025-11-17) surveys three benchmarks — METR's time…
- Task Time-Horizon Scaling×2
Gdpval Benchmark — the deflationary counterweight taught alongside this metric: win rate against a…
- Claude Opus 4.8
The GDPval-AA row is the v1 board as reported in 4.8's own system card. Artificial Analysis later…
- Context Advantage, Not Taste
Gdpval Benchmark — the asymmetry run as an experiment on paid work. GDPval writes all the context a…
- Cost-per-Task Over Cost-per-Token
Gdpval Benchmark — what the GDPval-AA row in the model-selection table above is built on, and where…
- Deep Research Agents
Read against GDPval from the same lecture, the two failures are one failure at different distances:…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated
Gdpval Benchmark — an instrument outside this page's four, and outside Steele & Cruz's seven,…
- Failures That Look Like Success
Gdpval Benchmark — the class as the dominant failure mode of an economically-framed benchmark…
- Jagged Intelligence (Ghosts, Not Animals)
Gdpval Benchmark — jaggedness at occupation granularity, and the reason its aggregate win rate…
- Measuring Beyond Accuracy Saturation
Gdpval Benchmark — the third response to "traditional benchmarks are saturating", after…
- Evals & Benchmarks
Gdpval Benchmark — OpenAI's late-2025 benchmark of real, economically valuable knowledge work:…
- Open Questions Backlog
Gdpval Benchmark ×2 (oldest 2d) — The paper itself is not in this corpus — every figure here is a…
- The Open-Weight Frontier Gap
Gdpval Benchmark — the benchmark under the GDPval-AA Elo board this page differences: real…
- OpenAI
Gdpval Benchmark — OpenAI's benchmark for whether a model's deliverable beats a 10-year…
- Returns to Expertise in Agentic Coding
Gdpval Benchmark — the benchmark that deliberately erases the premium: tasks are authored by…
Related articles
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Production-Sourced Evaluation
Building benchmarks from de-identified real production usage rather than synthetic or hand-authored tasks; DRACO's cent…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
