H
Howardism
Plate IIAI Economics & Labor中文HOWARDISM

Returns to Expertise in Agentic Coding

Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the actions and 5× the output per prompt, reach verified success ~2× as often, and abandon stuck sessions far less; every occupation lands within 7pp of software engineers; gains are concentrated novice→intermediate, with mastery adding little

Article metadata
Publication details
Published:June 17, 2026
Filed:Concept
Domain:AI Economics & Labor
Reading:37 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Returns to Expertise in Agentic Coding

Sources#

Summary#

The headline finding of Anthropic's economic-research report Agentic coding and persistent returns to expertise (Hitzig, Massenkoff, Lyubich, Heller, McCrory, June 2026): what amplifies an AI coding agent is the user's domain expertise, not their coding proficiency. Across ~400,000 Claude Code sessions, the more a person understands the problem they are solving, the more work the agent does per instruction, the more often the session succeeds, and the more readily it recovers from trouble. Coding background, by contrast, barely matters: in code-producing sessions every major occupation lands within seven percentage points of software engineers. The report's one-line thesis — "coding agents are not substituting for domain expertise; the more understanding a worker brings to an agent, the more quality work the agent is able to do" — is the empirical confirmation of "you can outsource your thinking but not your understanding".

Evidence note. empirical — measured from a privacy-preserving (Clio) analysis of ~400,000 interactive sessions from ~235,000 people, Oct 2025–Apr 2026, with classifiers (Claude Sonnet 4.6) validated against automatic telemetry and regressions with controls and confidence intervals. Two honest caveats keep it short of a clean external benchmark: it is first-party — Anthropic measuring its own product, on its own telemetry, with its own classifiers (not independently reproducible) — and it excludes headless (claude -p), SDK, and third-party-IDE usage, a "substantial share" of real activity. Outcomes are transcript-inferred proxies, not observed real-world results.

Expertise is task-specific, not a résumé#

The expertise rating (a five-point novice→expert scale) is not job title or general ability. "A senior engineer asking their first Rust question is a beginner at Rust. An accountant who has never used Python, but tells Claude exactly which reconciliation rules a script must enforce and catches the edge case it mishandles at month-end close, is an expert at that task." The classifier reads three signals:

  1. Precision of framing — how specifically the user directs the work.
  2. What they ask Claude to verify — experts specify the checks that define "done."
  3. Who corrects whom — does the user correct Claude, or does Claude correct the user.

This is exactly the understanding residue made measurable: the signals are the surface evidence of an internal model good enough to direct, judge, and verify.

The amplification: experts get more agent per prompt#

Expertise scales how much autonomous work each human prompt sets off (the action chain):

User expertiseActions per promptOutput per prompt
Novice~5~600 words
Expert~12~3,200 words

More than twice the actions, five times the output. The gap holds within every work mode and every task-value band, and survives a regression controlling for work mode, task value, month, occupation, and model family: +9% actions and +13% output per expertise level (p < 0.001 at each adjacent step). An expert isn't just luckier — each instruction they give safely unlocks a longer leash.

The success gradient#

The more expertise a session exhibits, the more likely it succeeds, on every measure. Success is defined two ways: judged success (a classifier reads the transcript and decides whether the person got what they set out to do) and the stricter verified success (judged-successful and at least one hard external signal — a matching commit/PR, passing tests, or explicit user affirmation).

ExpertiseVerified successAt least partial success
Novice15%77%
Intermediate–Expert28–33%91–92%

The curve is concave: most of the gain is novice→intermediate; intermediate→expert is modest. A working grasp of the domain captures most of the benefit; deep mastery adds only a little. (These are adjusted rates — comparing sessions of the same work mode, value band, month, subject, and occupation type.)

Two corollaries about recovery, which is where expertise earns its keep:

  • Recovering from trouble. Among sessions that "hit trouble" (verified failure signals — errors, failed tests, repeated attempts, user frustration), verified success rises from 4% (novice) to 15% (expert); partial success from 60% to 80–81%. Part of the value of expertise is the ability to steer a struggling agent back on course. (Caveat: experts hit trouble less often, so their troubled sessions are on harder problems — estimated task value of a troubled session roughly doubles from novice to expert — so some of the recovery gap reflects novices stuck on routine problems vs. experts stuck on genuinely hard ones.)
  • Abandonment. A troubled session is abandoned when it is judged failed and zero lines of code were written. 19% of novice sessions end abandoned, against 5–7% for everyone else. The least experienced give up when stuck.

Occupation matters less than expertise#

The complementary half of the story, and the strongest evidence for software democratization: a coding background is becoming less relevant to coding success. Occupation is inferred (the classifier is explicitly told not to treat the act of coding as evidence of a coding job — a lawyer scripting contract-clause checks is mapped to Legal, not Software).

  • Software-related occupations reach verified success ~30% of sessions overall; other professions ~26%. In code-producing sessions, 34% vs 29% — and partial success 89% vs 88%.
  • Every one of the ten largest occupations lands within seven points of software engineers on verified success in code-producing sessions; the five-point software/non-software gap has neither widened nor narrowed over seven months.
  • Management occupations edge out software engineers on verified success. The report's reading: management skills (delegating, specifying, confirming) transfer to directing an agent — "perhaps acting like a manager confers greater success." (Measurement caveat: verified success partly rests on explicit in-transcript confirmation, and managers may simply say when they got what they asked for.) This is the constructive counterpart to HBR's accountability critique — the skill of bounded delegation helps; the org-chart framing of agents-as-employees is what backfires.

What it means for the labor market#

The report frames itself as an early read on knowledge-work transitions. Two readings, in tension only superficially:

  • Substituting for coding skill. Implementation-heavy work that used to require a coding background is being absorbed; "a coding background [is becoming] less relevant to successful programming." This is the floor rising — "a person with command of a domain, in any field, may now be able to do technical work they previously could not."
  • Rewarding domain understanding. Simultaneously, the gains accrue to whoever brings the firmer grasp of the problem. "A person without any such expertise will get far less from the same tool." This is the residual human comparative advantage showing up in usage data: not coding, but knowing what to build and being able to verify it.

The report names the metric to watch: if the returns to expertise begin to decrease over time, that signals models are starting to supply the judgment users currently bring — i.e., taste becoming "just another capability". As of this data, the returns are persistent.

The premium, priced in vacancies. Indeed Hiring Lab (Gallacher, July 2026) finds the same shape on the hiring side rather than inside sessions. US software-development postings rose ~15% from Claude Code's February 2025 launch through mid-2026 while overall postings fell 7% — but the rebound is concentrated: 71% of the May 2025 → May 2026 increase came from senior roles and 37% from postings whose title mentions AI (the two overlap). Gallacher's reading is this page's thesis in an employer's words — "demand is growing for experienced professionals who can work with AI, not necessarily a broad-based recovery across all software roles." Two things keep it a corroborating signal rather than proof. It is Indeed analyzing its own job board in a blog post, so the sample is one platform's vacancy flow and the causal story (agentic tooling raised demand for expertise) is asserted from a coincidence of timing, not identified — see Firm AI-Spend Intensity and Headcount Growth for the full evidence note. And "senior role" is a title, not the task-specific expertise this study measures: the classifier here rates a senior engineer asking their first Rust question as a beginner, while a job posting cannot. Seniority-titled demand is the closest labor-market proxy for the expertise premium currently available, and it is a coarse one.

The composition check: experts using AI on their inexpert tasks#

Google ATLAS (July 2026) supplies a cross-lab observation that looks like a contradiction and resolves into a mechanism. Two facts hold simultaneously in its Gemini data:

  • The users skew expert. A 1% increase in an occupation's median earnings is associated with >2.5% higher AI usage intensity (1.86 controlling for education); weighting US median earnings by conversation volume moves it from $62,252 to $82,919.
  • The tasks skew inexpert. Sorting ~19,000 O*NET tasks into expertise quartiles by the Autor–Thompson word-rarity measure, usage is most over-represented on the lowest-expertise non-routine cognitive tasks — 2.6× baseline for Q1 against a flat 1.6–1.8× for Q2–Q4.

So the modal work interaction is a high-expertise worker pointing AI at the least expert-demanding parts of their job. That is this page's thesis observed from the outside: the human keeps the judgment and offloads what doesn't need it, which is why the expertise premium persists rather than dissolving. Autor & Thompson's model says which way it cuts — automating an occupation's inexpert supporting tasks raises the scarcity of the remaining human expertise (wages up, employment down), while automating its expert tasks erodes entry barriers and depresses wages. ATLAS's snapshot points at the first.

The caveats are real: ATLAS measures consumer surfaces and free API only (no enterprise, no agentic coding at scale), it is a two-week snapshot with no time dimension, and its expertise measure is a lexical proxy — task-statement word rarity — not the behavioral three-signal classifier this study uses. The two are measuring "expertise" on different objects: ATLAS rates the task, Anthropic rates the user's handling of it. See Task Saturation: Broad but Shallow AI Diffusion.

The same premium, in payroll instead of sessions (Canaries, August 2026)#

This study measures the expertise premium inside ~400K sessions over seven months. Brynjolfsson, Chandar & Chen (revised August 2026) measure something adjacent across millions of US workers over four years, on ADP administrative payroll through June 2026 — and the direction matches.

  • Complementarity accrues to the experienced, substitution to the young. Regressing occupation-level employment change (Nov 2022 → Jun 2026) jointly on standardized automation, complementarity and overall-usage exposure from the Anthropic Economic Index's automative/augmentative query split, by age band: the automation coefficient runs −0.098*** at 22–25 and shrinks monotonically to −0.006 at 50+, while the complementarity coefficient is ~0 for the young and turns positive and significant only at 41–49 (+0.024**) and 50+ (+0.015*) (Table 3, March 2025 AEI release; verified against pdftotext -layout, replicated on the pooled later releases).
  • And it attaches to tacit knowledge specifically. Laying a codified/tacit knowledge index over the same panel (Codified vs Tacit Knowledge Exposure), experienced workers in high-tacit occupations grow faster than experienced workers in low-tacit ones — and that gradient survives a college-share control, while the codified gradient for the young does not.

Why this is corroboration and not confirmation. The two studies measure different objects, and the difference is the same one this page already flags about Indeed's "senior role" titles. Anthropic's classifier rates task-specific expertise from behavior — a senior engineer asking their first Rust question scores as a beginner. Canaries has no user-level expertise measure at all: its proxy is age band × occupational knowledge type, which is tenure and job category, not command of the problem in hand. So payroll shows that experience is currently being complemented; it does not show that the behavioral thing this study measures is what is being paid for. What it adds is duration and scale — four years and millions of workers against seven months and 400K sessions — on a first-party-free instrument, which is exactly the external corroboration this study's own evidence note says it lacks.

And it puts a second, slower needle on the forward test. This page names the metric to watch: returns to expertise decreasing would signal models supplying the judgment users currently bring. The payroll analogue is the complementarity coefficient for 41–49 and 50+ narrowing toward zero. Both currently point the same way: persistent.

An expert population that was measured and not cut by expertise (September 2026)#

Worth recording as a negative result about instruments, not as evidence for the thesis. Google/DeepMind's science follow-on to ATLAS (AI Adoption in Scientific Work, empirical with vendor COI) surveys 637 working scientists — 356 in senior roles (PI, professor, lab director, industry R&D manager), 234 mid-career, 47 early-career, mean just below 13 years of research experience — and reports that just under three quarters save time on net, averaging 6.9 hours per week, with 46.6% using AI daily.

That is a large self-reported dividend on about as expert a population as the corpus contains, and it says almost nothing about returns to expertise, for two reasons. First, it is self-report with the sponsor's interest pointing one way, from a screened non-probability panel whose authors name selection toward AI-enthusiastic respondents as a reason the savings may be overstated. Second, and more instructive: the survey holds career stage for every respondent and publishes no cut by it — not adoption, not hours saved, not the share of saved time lost to verifying output. The gradient this page exists to measure was available in that dataset and was not run. The one gradient the report does publish is on the task, not the person: its ~360K science interactions score +26% on ATLAS's domain-expertise classifier against the average work conversation, which is the same lexical-proxy measure flagged above, applied to a population rather than to a quartile of tasks.

The novice side, randomized — and its scope condition (2026-08)#

This page measures the gradient from the expert end: understanding amplifies, and the curve is concave, so novice→intermediate captures most of the gain. A preregistered 2×2 RCT supplies the missing novice-end datum under randomization rather than telemetry — and it is larger than the telemetry implies. Training novices to think, or giving them LLMs? Evidence from an RCT (Bocconi × OpenAI, CEPR DP21882, empirical, n=1,053 first-year undergraduates) gives half the sample ChatGPT Edu (GPT-4o) for a 45-minute merchandising case and scores the output 1–5 on two codified marketing criteria via three of twenty trained, condition-blind master's raters. LLM access raises the expert-graded score by +0.862 on a control-group estimate of 2.09 (s.e. 0.079, p<0.01) — about 1.0 SD — and moves solutions closer to three independently interviewed domain experts (+0.026 / +0.028 / +0.044 cosine similarity, all p<0.01). Post-double-selection lasso strips coherence, idea count, diversity and every text-style feature and ~42% of the effect survives: the model is improving substance, not only polish. Novices with no domain understanding at all land measurably nearer the expert answer.

The finding that bites this page hardest is the other arm. The same experiment randomizes a taught cognitive skill — a ~6-minute causal-reasoning game — and it works on reasoning (mechanism identification ~0.55 SD, falsification logic ~0.85 SD, idea diversity +0.47 SD between solutions, none of it crowded out when a model is in reach). It does nothing for the evaluated score; the coefficient is negative (−0.111, p<0.01). On this task, the human's understanding was not the amplifier — the model was.

The reconciliation is the scope condition, and both papers state it. Anthropic measures open-ended agentic work where the problem must be framed, the checks specified, and the model corrected; the RCT deliberately picks a problem where "LLMs are likely to be well-trained" — a documented marketing framework, a codified rubric, a single 45-minute deliverable, and a population that has no domain expertise to contribute. That is the corner of the space where the expertise premium should be smallest, and it is: the LLM supplies what expertise would have. Read together they bound the claim rather than contradict it — the returns to expertise are a property of the task's definition, not of the user alone, and the concave curve's novice end is steepest exactly where the problem is well-specified enough that the model already knows the answer. The RCT's own warning about that corner is the one this page should carry forward: what the rubric rewarded was the model's coherent, idea-dense conformity, while the trained human's distinctiveness was scored down (between-solution diversity loads at −7.676, ≈ −0.29 points per SD).

Where the expertise is supposed to come from (Martin, 2026-08)#

This page measures that expertise amplifies agents and that the gains concentrate novice→intermediate. It does not say how a novice gets to intermediate once agents have absorbed the tactical work — the pipeline question. Robert C. Martin gives the corpus's most concrete proposal, and is candid that it is a guess: "I get this question a lot. I don't have perfect answers to it because I really don't know" (Uncle Bob on Software Fundamentals in the Age of AI, 2026-08-19, practitioner-opinion).

His three-stage path:

  1. Write code, for about a year. "So that you know what the agents are dealing with." He generalises an older piece of his own advice — spend a weekend writing assembly so you know what is under the abstraction — into a ladder: binary, assembly, C, Python, then agent work.
  2. Be treated as an agent. "The young person just coming out of training should be treated like an agent. The lead engineer who's got a bunch of agents running and he's being strategic — he should look at you as an agent, give you the same kind of tasks that the agents have, and subject you to the same deterministic tools that the agents have to use. And you should spend several months in that state being horribly unproductive but learning a hell of a lot."
  3. Then run agents. "By the time you've gone through that gauntlet, maybe you can be trusted to run an agent of your own."

What makes this more than a training anecdote is the specific skill he says the gauntlet teaches, which is not code quality. Asked how he knew his agents were failing in December 2025, he separates two signals: "Early on I was just looking at the code and seeing the dog doo. That wasn't the important part. The important part was the next step where I watched them thrash. I could see the agent struggle and I recognized the struggle since I have been through that struggle." The transferable asset is pattern recognition for being stuck, acquired by having been stuck. And he names the failure directly: "one of the issues is the novice would come in and not recognize the struggle" (Agentic Technical Debt for what the struggle indicates).

Three problems the proposal carries, two of which Matt Pocock raises on the spot:

  • The economics. "If you have someone who's just sitting on your books doing tactical work, surely as a company you're going to look at that and go, why are we hiring this person when we've got a swarm of five hardeners that can beat it at a fraction of the cost?" Martin has no answer. Several months of deliberately unproductive labor is a cost someone must absorb, and this page's own finding — that gains concentrate novice→intermediate — makes the investment rational for the worker without making it rational for any particular employer.
  • The feedback loop was always the problem. Pocock's other point is that strategic judgment traditionally took years to develop because mistakes surfaced nine months later, and people changed jobs first. He suggests agents compress that loop, which if true is a reason to expect faster expertise formation, not slower — and cuts against the apprenticeship's necessity.
  • It inverts Martin's own principle. He argues elsewhere that human disciplines should not be imposed on agents because they are adaptations to human cognition (Impose Values, Not Disciplines); here he imposes the agent's working conditions on a human. He does not address the asymmetry.

His fallback is reading: Tom DeMarco, Ed Yourdon, The Pragmatic Programmer — "the old books, the ones that nobody reads because they're old… you'll have to filter out some of the archaic stuff, but that's when these lessons were learned."

Connections#

  • Robert C. Martin (Uncle Bob) — the apprenticeship proposal, the skill it is supposed to teach, and the economics nobody has answered

  • Impose Values, Not Disciplines — the principle his own proposal runs in reverse

  • Task Time-Horizon Scaling — the expertise premium measured as a speed ratio and then handed to the model: on internal pull requests, contractors without codebase context are 5–18× slower than maintainers, and (per CS329A lecture 8) model performance tracks contractor times. The premium this page attributes to domain expertise is the same quantity METR's horizon is missing

  • GDPval Benchmark — the benchmark that deliberately erases the premium, and since 2026-09-10 with the primary paper's expert bar rather than a lecture's. Tasks are authored by practising professionals — minimum 4 years in the occupation, 14 years on average, screened for promotion and management history through a resume review, video interview, background check, training and a quiz, with fewer than 10% of applicants accepted (the wiki previously carried "10+ years", read off a slide) — and each task is built around work product the expert actually produced, with all their tacit context written into the prompt. The model is therefore graded against an expert on the one axis where the expert's advantage has been handed over. Restore the ambiguity by cutting prompts to 42% of their length and GPT-5-high's win-or-tie rate falls 47.7% → 44.3% while the qualitative failure is larger than that: the model stops knowing what to work on

  • Agent Review Comment Resolution — the same premium on the receiving end of review. Across 341 repos, core developers (top 20% by authored-plus-reviewed PRs, computed per repository) resolve 78.1% of Copilot's agent review comments and 70-80% of every comment category — with their share highest on solution approach, documentation, naming and code organization, and lowest on functional defects (29.5% peripheral). Acting on an agent's design feedback takes project knowledge; fixing a defect it found does not. The mirror finding is the sharper one: rejecting a wrong agent suggestion is the one behaviour that leans peripheral (35 vs 32 cases), so the expertise premium is on knowing what the project intended, not on catching the model out

  • Controlled Variance: AI's Edge as Reduced Dispersion — the counter-case, and it cuts two ways. In a randomized hiring experiment, 131 experienced recruiters forecast the effect's direction wrong on the task they were expert in (36% expected the AI arm to receive lower offer rates, 48% lower retention, 61% lower interview quality; the AI arm got 12% more offers). Expertise did not confer foresight about its own automation. But the paper's own reading preserves this page's thesis on the other margin: interviewing automated cleanly while evaluation — the judgment half — stayed human and became the process bottleneck, "redirecting recruiter expertise toward evaluation"

  • The Solo-Authorship Rebound — a levelling signature on a different outcome, and not a contradiction. Across 300M+ OpenAlex works the post-2022 switch into solo authorship is largest among the least prolific authors (Δβ +1.16 for 1–5 lifetime works vs +0.21 for 6–20) and only weakly largest among the most senior. That is AI cutting the fixed cost of producing output at all, sitting beside this page's amplification of output per prompt — and neither measures whether the low-output authors' solo papers are any good

  • The Tragedy of the Cognitive Commons — the regeneration question this page's cross-section cannot see: expertise measurably amplifies an agent today, and Lovett argues the entry-level work that builds that expertise is what AI removes first

  • Task Crossover — the unreconciled tension: expertise is what amplifies an agent here, yet crossover finds a large share of AI work happening outside the user's expertise. Both rest on usage telemetry, and nothing yet says whether borrowed work is done as well

  • Implementation Abundance Inverts Product Work — the product-process face of "judgment outlasts cheap execution": as implementation cheapens, curation/taste becomes the expensive step

  • Task Saturation: Broad but Shallow AI Diffusion — the cross-lab composition check: ATLAS finds high-earning workers over-using AI on the lowest-expertise cognitive tasks, which is this page's mechanism (offload the inexpert, keep the judgment) seen in Google's usage data rather than Anthropic's

  • Role Averaging, Not Role Elimination — the empirical backbone of "specialties don't disappear": domain expertise still decides success even as roles average

  • Outsource Your Thinking, Not Your Understanding — this is the empirical proof of Karpathy's thesis: success tracks understanding of the problem, not the ability to type code; the non-delegable residue, now measured

  • Printing Press Software Democratization — Cherny's "the best person to write accounting software is a good accountant, because coding is the easy part" is exactly the every-occupation-within-7pp finding; this is the hard data the analogy was waiting for

  • Vibe Coding vs. Agentic Engineering — "floor up, ceiling held": occupation-doesn't-matter is the floor rising; expertise-still-decides is the bar that stays

  • Research Taste as the Human Bottleneck — the "if returns to expertise decrease, the model is supplying judgment" test is the labor-data version of "is taste a durable moat or the next jagged valley?"

  • Planning / Execution Division of Labor — the mechanism of amplification: expert framing safely lengthens the action chain Claude runs per prompt (5→12 actions)

  • Agentic Coding Work-Composition Shift — the companion finding from the same study: what the work is and how it shifts over the seven months

  • AI Employee Framing — managers' edge here (delegation skill transfers) is the constructive flip side of HBR's warning (employee framing diffuses accountability); skill helps, org-chart symbolism hurts

  • Verification as the New Bottleneck — "what they ask Claude to verify" is one of the three expertise signals; the expert is the one who can specify and check, which is exactly the bottleneck role

  • Engineer PM Convergence — domain/product understanding as the bottleneck skill, seen in the success data

  • Jagged Intelligence (Ghosts, Not Animals) — experts recover from the agent's spiky failures; novices abandon — staying in the loop pays measurable dividends

  • Claude Code — the product the entire study measures

  • METR — the report cites METR's time-horizon ceiling as the capability frontier this usage sits below

  • Conversation-to-Delegation Shift — OpenAI's Codex study cites this report (Hitzig et al. 2026) and reaches the same conclusion from usage data: as work becomes delegation, the binding skill is domain understanding + supervision, not execution

  • Organizational Complements to AI — the cited Hitzig et al. argument restated as economics: supervision/verification/coordination and domain expertise are the binding complements that gate AI's value

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the AEI Cadences survey confirms this from the worker's mouth: 15+-year workers report AI can do ~10pp less of their work, naming judgment, context, and relational work as what AI can't touch — tacit expertise as the residual. That page also carries the augmentation-pays finding: ranking 872 occupations by share of augmented vs. automated Claude queries, jobs in the ninth augmentation decile out-earn jobs in the ninth automation decile by about $7,000/yr (Steele & Cruz 2026). The correlational, 2022-salary, work-and-personal-use-mixed version of this page's thesis — pay tracks the work that is hard to hand over

  • The Automation–Optimism Link — the mirror gradient: experienced workers are more skeptical of AI's reach, yet heavy delegators are the most optimistic — expertise and enthusiasm pull in opposite directions

  • Experimental Learning Impact of Generative AI — the learning-task echo: AI's gains skew to higher-ability students and to augmentation use (AI deepening understanding), the same "benefit accrues to whoever brings/builds understanding" pattern this study measures in agentic coding — and, on the same page, the Bocconi RCT that supplies this page's novice-end bound and its scope condition (a well-defined problem on which the model is already well-trained is where the expertise premium is smallest)

  • Conversation Artifacts — "the human stays involved in high-value work" (more turns and more Claude output as wages rise) is the augmentation reading of expertise amplifying the agent

  • AI Usage Cadences — off-hours Claude work skewing to higher-wage occupations is consistent with expertise-heavy work being where AI lands first

  • Anthropic Economic Index — the research program this study belongs to; Cadences is its next report

  • Context Advantage, Not Taste — the instrument. If the human's role is a closable information asymmetry rather than a faculty, the measured decline in the expertise premium is that gap closing, tracked in usage data; the two framings disagree about what the decline means

  • Unknowns as the Agentic Bottleneck — the mechanism behind the expertise premium: Thariq Shihipar observes that "the best agentic coders have relatively few unknowns" — the expert's map already matches the territory, and they assume unknowns rather than believing the map complete

  • Review as the Control Point — the same claim on the review side: reviewer expertise + disposition is the first of the three moderators that decide whether a coding agent helps or harms software. Expertise amplifies (and protects) at the review keyboard as it does at the authoring one — the CMU theory is the mechanism story for why this study's expertise premium should persist

  • Market-Priced AI Exposure (the AI Premium) — the asset-pricing echo: the AI premium loads on the intensive margin (paid/core and seasoned users, long prompts price AI risk; casual/new use does not), and the market-implied skill map rewards interactive/relational/communication work and penalizes analytical/scientific — the market's version of this page's "depth of understanding amplifies the agent" and the survey's "experienced workers name relational judgment as the residual AI can't touch"

  • AI-Native Organization — the practitioner-side amplification claim to weigh against this data: Tan's self-measured ~400x (self-deflated to 8x–80x) vs. the measured 2× actions / 5× output premium; his "2x people and 100x people use the exact same Claude" attributes the spread to wiring, not expertise — a complements story this study doesn't test

  • Firm AI-Spend Intensity and Headcount Growth — the two labor-demand instruments that bracket this page's premium, and they disagree on composition: Indeed's postings rebound is 71% senior (the expertise premium priced in vacancies), while Ramp's spend-linked firm panel finds entry-level headcount growing fastest (+12.0%) at intensive adopters. Stock vs. flow, adopting firms vs. one job board — reconcilable, unreconciled

  • Owning Your Externalized Cognition — the premium this page measures, proposed as portable property: if domain judgment is what amplifies an agent, writing it down as skill files turns it into a transferable asset. This page's concave curve complicates that — gains concentrate novice→intermediate and mastery adds little, so the library of a true expert may be worth less as career capital than the doctrine assumes

  • AI and Market Power — the firm-level analogue of this page's individual premium, with the same double edge: OECD find the benefits of GenAI "unlocked more easily by firms that have higher productivity and capabilities," measured as a monotone gradient across GenAI-exposure quintiles in productivity, markups and tertiary-educated workforce share (0.04 → 0.46). Where this page finds expertise amplifying an agent within a session, that one finds firm capability amplifying adoption across a whole economy — and draws the same worry, that the capable gain most and the gaps widen

  • Systems Thinking Over Specialization — the hiring-side reading of both halves: Elizabeth Stone's "great engineering is scarce" is the persistent premium, and her "specialists can learn that quickly now" bets on exactly this study's concave curve (working grasp is cheap to acquire; mastery isn't)

  • Codified vs Tacit Knowledge Exposure — this page's thesis at population scale and on a different instrument: experienced workers in tacit-heavy occupations grow faster in four years of ADP payroll, and the complementarity coefficient from the Anthropic Economic Index usage split is positive only for the 41+ bands. The gap worth holding onto is that payroll measures tenure, not the task-specific expertise this study rates from behavior

  • Cognitive Capability Profiling for Task Suitability — the complement this page names, excluded by construction from the other side. Prunty et al. build their capability set from "core cognitive processes relevant to general workplace activity rather than the specialist skills that differentiate one worker from another," and their 410 respondents rate the importance of those core capabilities for their activities — so domain expertise, the variable that most amplifies an agent in the 400K-session study, is deliberately outside the instrument. The consequence shows up in their own results: workplace activities converge on a shared cognitive core and the suitability ranking of six AI systems barely changes across 18 activities, which is what you would expect from a measure that has removed the differentiating variable. Their requirements side also runs into the tacit-knowledge problem from the measurement direction — capabilities like metacognition and spatial reasoning are nominated rarely, the authors suspect because automatic processes are not salient to conscious reflection, not because they contribute little

Open Questions#

  • The forward test the report itself names: do the returns to expertise persist, narrow, or invert as models improve? A decrease would mean models are absorbing the judgment users currently supply.
  • Outcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the whole success gradient — is real outcome vs. who-narrates-success-in-the-transcript?
  • The study excludes headless / SDK / IDE usage (a "substantial share"). Does the returns-to-expertise pattern hold in non-interactive and pipeline use, where there is no human steering mid-session at all?
  • Is "intermediate captures most of the benefit" stable, or an artifact of current model capability — i.e., will the concave curve flatten further (everyone converges) or steepen (mastery starts to separate again) as models get better?

Derived#

Sources#

  • GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan et al. (19 authors, OpenAI), arXiv 2510.04374 v1, 2025-10-05, 29pp (empirical). Cited here for §2.2's expert-recruitment bar (4-year floor, 14-year mean, screening pipeline, <10% acceptance, ≥5 professionals per occupation) — which corrects the "10+ years" figure this page had carried from a lecture transcript — and A.2.7's under-contextualization arm. Full treatment and COI on GDPval Benchmark
  • Agentic coding and persistent returns to expertise — Anthropic Economic Research, June 2026; §"The returns to expertise", §"Occupation may matter less than expertise", §"Looking ahead"
  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, Voice AI in Firms (arXiv 2607.28222, 2026-07-30; empirical, pre-registered RCT): §3 intro (the recruiter forecasts), fn. 5 and §8 (recruiter expertise redirected toward evaluation), §7.1 (the evaluation stage as the new bottleneck). Parse warnings and full treatment at Controlled Variance: AI's Edge as Reduced Dispersion.
  • AI and Job Postings: From Destruction to Creation? — Guillermo Gallacher, AI and Job Postings: From Destruction to Creation? (Indeed Hiring Lab, 2026-07-08; empirical administrative postings data, but a blog post by the platform whose data it is). Key points list and §"A senior, AI-fluent rebound" — the 15% / 7% / 71% / 37% figures, all stated in prose as well as in chart captions. The international comparison's country identities and both scatter plots' correlation statistics are chart-image-only and are cited nowhere in this vault. Full evidence note and the contradicting entry-level result at Firm AI-Spend Intensity and Headcount Growth.
  • Helping People Choose Careers in the Age of AI — Steele & Cruz, arXiv 2607.15506 (2026-07-16), empirical; §5.6 + Fig. 17 (mean salary by augmented vs. automated Claude usage decile). The $7,000 ninth-decile gap is stated in prose; the authors attach two caveats (2022 salary data, and the September 2025 AEI release cannot separate work from personal use). Parse warnings for this source are recorded on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated.
  • Training novices to think, or giving them LLMs? Evidence from an RCT — Asirvatham et al. (12 authors, Bocconi University × OpenAI), CEPR DP21882, August 2026, empirical (preregistered 2×2 RCT, n=1,053, class-level randomization across 13 sections). Cited here for §4.2 (awareness/usage score and expert similarity) and §4.4 (PDS-lasso decomposition). COI: three authors are OpenAI-affiliated or OpenAI-contracted, ChatGPT Edu is the treatment, and the similarity and idea measures run on OpenAI models — the human-rated score is the one outcome outside that loop, and it carries the largest pro-LLM effect. Parse warning: PDF-derived, with an unflagged standard-error row-shift across Tables A.7–A.14 (Table A.11 misattributes the 0.044*** Expert-Alumnus GPT coefficient to the Causal row; Table A.13 drops two coefficients outright); all figures here were recovered from pdftotext -layout. Full treatment and the causal-training arm at Experimental Learning Impact of Generative AI
  • Uncle Bob on Software Fundamentals in the Age of AI — Robert C. Martin with Matt Pocock, 2026-08-19 (practitioner-opinion; auto-caption transcript): the write-code / be-an-agent / run-agents pipeline, recognizing the struggle as the transferable asset, and Pocock's unanswered objection about who pays for it
  • Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence — Brynjolfsson, Chandar & Chen, Canaries in the Coal Mine?, Stanford Digital Economy Lab, revised August 2026, 140pp (empirical, ADP payroll microdata through June 2026). Cited here for Table 3 and §2.5 (the automation/complementarity coefficients by age) and §2.5 + Online Appendix J (the codified/tacit gradient for experienced workers). Used as population-scale corroboration of the expertise premium, with the explicit caveat that its expertise proxy is age × occupation rather than behavior. Full treatment on Codified vs Tacit Knowledge Exposure
  • Google AI & Economy ATLAS: AI in Science (September 2026) — AI in Science: Early Insights (Google, Google DeepMind, MIT FutureTech, September 2026, 42pp; empirical with vendor COI). Cited here only for Appendix 3 (sample composition and seniority bands), §5.2 (6.9 hours/week self-reported net saving, 46.6% daily use) and §3.1 (the +26% domain-expertise score on science interactions) — as an expert population measured without an expertise cut. Full treatment on AI Adoption in Scientific Work
§ end
Cited by 52
Related articles