Sources#
- “We Must Act Now”: Sixteen Nobel Laureates Join Leading Economists and AI Researchers in Call to Prepare for AI’s Economic Transformation
- AI and Job Postings: From Destruction to Creation?
- Anthropic Economic Index report: Cadences
- Helping People Choose Careers in the Age of AI
- The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
- Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
Summary#
"AI exposure" is used loosely to mean very different things. The Anthropic Economic Index's Cadences survey (June 2026) makes the distinctions sharp by putting four measures side by side — two derived from behavior, two from what workers say — and showing how they order and diverge. The upshot: what AI is observed doing, what it could theoretically do, and what workers believe it can do are three different numbers, and separating them is a prerequisite for reasoning about AI's labor impact.
Evidence note.
empirical— observed/theoretical exposure from prior AEI work and standard task-based measures; reported/anticipated from the linked survey (self-report, non-representative sample, binned responses coded at bin midpoints). See Anthropic Economic Index.
The four measures#
| Measure | Definition | Source |
|---|---|---|
| Observed exposure | Share of an occupation's tasks already seen being done with Claude | Usage telemetry (prior AEI reports) |
| Theoretical exposure | Share of tasks an LLM could theoretically do — an upper bound | Task-based external measure |
| Reported exposure | What share of their work tasks respondents say AI can do today | Survey |
| Anticipated exposure | What share they expect AI to handle in 12 months | Survey |
Ordering and correlation. Reported exposure is positively correlated with both observed and theoretical exposure — people's beliefs track reality. But the levels diverge systematically: reported > observed (the survey over-reaches heavy users, who see more of what AI can do), and theoretical > reported (theoretical is an upper bound on the possible, not a measure of current use). ~6 in 10 respondents pick a higher band for next year than today; over ⅓ expect AI to do most or nearly all their work tasks within a year.
The rising tide: anticipated growth is uniform#
The striking survey result: while current perceptions vary with who and where you are, expectations of future progress are strikingly uniform. Plotting anticipated exposure against observed or theoretical exposure, the best-fit lines are roughly parallel — a software engineer and a construction manager anticipate about the same increment of progress in their own field over the next year, regardless of how exposed they already are. Everyone expects the tide to rise by a similar amount. (Caveat: midpoint coding of binned responses biases the slopes toward zero, so the parallel lines are read qualitatively.)
The cross-sectional gradients#
Reported/anticipated exposure vary systematically along three axes:
- ↓ with country GDP — reported exposure is ~10pp lower in high-income countries. Consistent with AI substituting for a larger share of lower-income daily tasks, even though occupation-level metrics run higher in advanced economies. The report ties this to the IMF complements argument: lower-income workers may lack the complementary skills/infrastructure that let AI augment rather than replace — and earlier AEI work found lower-income economies use Claude in more automated ways.
- ↓ with experience — workers with 15+ years put the share ~10pp lower than first-year workers. In follow-ups, experienced respondents pointed to judgment, contextual awareness, situational reasoning, and relational/interpersonal work (building trust, managing people) as things AI cannot replicate. This is Returns to Expertise in Agentic Coding read from the survey side: tacit, context-specific expertise is what the experienced believe AI cannot touch.
- ↑ with automation share — heavier delegators report and anticipate higher exposure (see The Automation–Optimism Link); delegation is informative about capability, and/or believers delegate more.
Notably, future-progress expectations are essentially uncorrelated with GDP and experience — the gradients are about today's perceived capability, not the rate of expected change (the rising tide again).
Why it matters#
Conflating these four numbers is how "AI can do X% of jobs" claims go wrong. Observed exposure is a floor (what's happening now), theoretical a ceiling (what's possible), reported the workforce's felt reality, anticipated its forecast. The gaps between them — reported exceeding observed, everyone expecting a uniform jump — are the interesting quantities, and each responds to different levers (capability, complements, belief).
Observed exposure, measured twice#
Google ATLAS (July 2026) supplies a second lab's observed exposure on the same O*NET substrate, and the comparison is a warning about treating any of these four numbers as a property of the economy rather than of the instrument:
| AEI | ATLAS | |
|---|---|---|
| Share of O*NET tasks with observed usage | 36% (Handa 2025) / 49% combined (Appel 2026) | ~20% |
| Automation share | 43–45% | <10% for non-routine cognitive |
ATLAS attributes the task gap to its stricter privacy thresholds (a task counts only above 25 unique users) and the automation gap to definitions — a five-category intent classifier where everything short of end-to-end execution reads as collaboration, against the AEI's binary split. Neither is wrong; they are different cuts of the same construct, which means observed exposure is a measurement convention, not a fact. ATLAS's Task Saturation: Broad but Shallow AI Diffusion adds a genuinely new cut to this taxonomy — not "what share of an occupation's tasks could or are exposed" but how deeply AI has penetrated the occupations it has reached at all (68% of occupations, 21% of tasks at the median). Breadth and depth are separable, and conflating them is how "68% of occupations" gets read as "68% of jobs at risk."
The error bar under all of it is now measurable in at least one case: see Usage-Telemetry Classifier Validation.
Seven instruments, head to head#
Steele & Cruz (arXiv 2607.15506, July 2026) run the experiment the section above only got to N=2 on: put seven occupational AI-exposure instruments — nine years of them — on the same O*NET occupations, standardize each to mean 0 / SD 1, and ask whether they rank jobs the same way. They do not. At N=7 the disagreement is close to total, and the paper's practical conclusion is that career advice built on any single instrument is advice built on that instrument's assumptions.
The seven#
(Tables 1 and 2 of the paper, both verified cell-for-cell at ingest — see Sources.)
| Instrument | How exposure is measured | What the number means | Distribution across occupations |
|---|---|---|---|
| Steele & Cruz 2026 | Anthropic + OpenAI queries | % of a job's tasks feasibly automatable | n=872, mean 23.6, SD 6.18, range 15.8–49.8 |
| Massenkoff & McCrory 2026 | Anthropic queries × Eloundou feasibility | share of LLM-feasible tasks Claude actually automates | n=872, mean 0.08, SD 0.12, range 0–0.73 |
| Eloundou et al. 2024 | GPT-4 task ratings, human-validated | share of a job's tasks an LLM can fully automate | n=872, mean 0.32, SD 0.19, range 0–0.84 |
| Felten et al. 2021 | crowd-sourced ability ratings (2,000 MTurk) | standardized suitability of the job's abilities for AI | n=759, mean 0.04, SD 1.0, range −2.67–1.53 |
| Webb 2020 | text-mining of AI patent titles | % of AI-patent verb-object pairs appearing in the job's tasks | n=719, mean 0.42, SD 0.29, range 0–1.49 |
| Brynjolfsson et al. 2018 | crowd-sourced DWAs against a 23-factor rubric | suitability for machine learning (SML), 1–5 scale | n=872, mean 3.46, SD 0.11, range 2.83–3.91 |
| Frey & Osborne 2017 | 70 hand-coded jobs → ML extrapolation to 702 | probability the whole job can be automated | n=689, mean 0.50, SD 0.38, range 0.003–0.99 |
Three things are visible before any correlation is computed. The estimands are not the same object: five score a basket of tasks, Felten scores abilities, and Frey is the only one that (in the paper's words) "classifies jobs as unitary entities rather than baskets of tasks, activities, or abilities." The distributional shapes have nothing in common — Frey is nearly uniform over [0,1], Massenkoff piles up at zero, and Brynjolfsson's SML uses 27% of its nominal 1–5 scale (every occupation in the economy lands between 2.83 and 3.91). And coverage differs by ~180 occupations (689–872), because the papers use different SOC vintages.
They cluster by data source, not by construct#
The scatterplot matrix (Fig. 6) shows no general positive correlation. Two pairs stand out, and both are explained by shared inputs rather than by convergent measurement:
- Steele & Cruz ↔ Massenkoff & McCrory: ρ = 0.89 — both built on 2025 Anthropic Claude usage.
- Eloundou ↔ Felten — the strongest of the rest, despite differing in level of analysis (19,000 tasks vs. 52 abilities), rater (GPT-4 vs. 2,000 MTurk humans), and question asked. What they share is a focus on theoretical generative-AI capability.
Everything older — Webb's patents, Brynjolfsson's ML rubric, Frey's bottlenecks — shows "very little correspondence with each other or with later measures." Instruments from Felten 2021 onward agree with each other somewhat more than the older ones do.
The generalization is uncomfortable for the four-way taxonomy at the top of this page: exposure instruments correlate when they share a data source, not when they claim to measure the same construct. Two observed-exposure measures agree because both read Claude logs; two theoretical-exposure measures agree because both ask what an LLM could do. Across families, agreement collapses. This reframes Market-Priced AI Exposure (the AI Premium)'s orthogonality finding (its market-implied map explains <2% of the variance in every task-based measure): the task-based measures are also largely orthogonal to each other, so orthogonality was never evidence that the market measure is the odd one out.
The punchline: twelve slots, eleven occupations#
Asked which occupations are most exposed, the six instruments the paper carries into its downstream analyses nominate almost entirely disjoint jobs:
| Instrument | Most-exposed #1 | #2 |
|---|---|---|
| Steele & Cruz 2026 | Database Warehouse Specialist | Business Intelligence Analyst |
| Eloundou 2024 | Telephone Operator | Telemarketer |
| Felten 2021 | Genetic Counselor | Financial Examiner |
| Webb 2020 | Wastewater Treatment Operator | Civil Engineer Technician |
| Brynjolfsson 2018 | Mechanical Drafter | Mortician |
| Frey 2017 | Telemarketer | Insurance Underwriter |
Twelve slots, eleven distinct occupations. The single overlap — telemarketer — is shared by Eloundou and Frey, two instruments whose projections are otherwise uncorrelated in Fig. 6, which the authors note as the coincidence it is. The paper attributes the spread to measurement philosophy: instruments built on generative-AI capability (Steele, Eloundou, Felten) surface technical and analytical roles; Webb's patent mining surfaces engineering; Frey's bottleneck model surfaces formulaic scripted work. If a worker asked any one of these six "am I the most exposed job in America," five of the six would say no.
Parse warning. This table is reconstructed. In the docling parse, Table 5's six model names collapsed into a single row-1 cell with rows 2–6 left blank, while the two occupation columns stayed in order — the classic collapse failure. The mapping above was recovered and verified against the PDF at ingest; it is not the parsed table.
The exposure–salary gradient changes sign by instrument vintage#
The four newest instruments show a positive, roughly linear relationship between standardized exposure and 2022 median salary, and the same against occupational complexity (an importance-weighted GWA sophistication score, range 18.6–74.5). Felten's is the steepest — its method ascribes high exposure to verbal and explanatory abilities, which puts attorneys and educators near the top.
The two oldest do not. Frey & Osborne 2017 runs strongly negative against salary; Brynjolfsson 2018 is flat against salary and slightly negative against complexity. Both have a stated cause rather than a mystery: Frey partially extrapolates from prior automation waves (routine manual and clerical work at the bottom of the wage distribution), and Brynjolfsson's rubric explicitly scores high-stakes tasks as less suitable for machine learning, which penalizes complexity by construction.
So "does AI land hardest on the top or the bottom of the wage distribution" — the question the entire white-collar-displacement debate turns on — is answered by the instrument's assumptions, not by the underlying economy. This is the strongest form of this page's thesis: the instrument sets not only the level of exposure but the sign of its headline distributional claim.
…and the sign also flips by time window, holding the instrument fixed#
Steele & Cruz vary the instrument. Indeed Hiring Lab (Gallacher, July 2026) holds one instrument fixed — the GenAI occupational-exposure measure from Hiring Lab's own AI at Work Report 2025, which the post links as its exposure source — and varies only the window over which the outcome is measured. It gets a sign flip out of that too:
- May 2022 → May 2026: the more AI-exposed an occupation, the more its US job postings fell. (Matching the New York Fed's finding that the AI-exposed decline in vacancies began before ChatGPT's release, which is itself a caution about reading the whole 2022–2026 slope as an AI effect.)
- May 2025 → May 2026: the relationship inverts — the more exposed, the more postings rebounded, led by software development.
Both are described as statistically significant and both are explicitly uncontrolled; the correlation coefficients and p-values live only in the post's chart images and are not quoted here. Two consequences for this page. First, "exposure" as used in the displacement debate silently bundles a direction ("exposed ⇒ shrinking") that the data supply only within a chosen window — an occupation can top the exposure ranking and top the growth ranking eighteen months later, so the exposure score orders occupations without signing what happens to them. Second, this is an outcome measured against exposure, not another exposure instrument: it belongs in the same family as the labor-market validation the head-to-head above never got — scoring the instruments against realized employment changes — except that its answer depends on which realized period you score against — a complication for the "score all seven against subsequent employment changes" test proposed in the Open Questions below.
And the theory side has the same indeterminacy, stated and then dropped. Banerjee & Singh's HAT model (arXiv 2607.20781, practitioner-opinion — a formal model with no data) derives a skill threshold μ* separating a human "protection zone" from a "vulnerability zone," and offers as its prediction P5 that "counterintuitively, within flat-Δ$' occupations, high-skill workers may be more vulnerable than mid-skill workers." But the corollary that generates μ* concludes the opposite of a determinate sign: "Whether highly skilled workers are protected or vulnerable is therefore task-dependent and cannot be resolved without domain-specific parameters." The threshold moves with baseline costs, organizational depth, AI capability, and the risk differential — and the paper says depth's effect on it is itself parameter-dependent. So a formal model and a seven-instrument empirical head-to-head reach the same place from opposite directions: the sign of the skill-vulnerability relation is not currently determined by anything, and the confident version of it — in either direction — is the assumption talking. What the model adds that the instruments cannot is a reason the relation could run either way: human compensation is convex in skill while AI cost is near-flat in capability, so which side wins depends on where the two curves cross, not on which agent is better at the job.
The seventh instrument: telemetry rebuilt as a projection#
Where the older six are projections of what AI could do (per raters, patents, or rubrics), Steele & Cruz build theirs from what people actually sent to models in 2025 — the first entry in this taxonomy to turn usage telemetry into a forward-looking exposure estimate rather than an observed-exposure count. See Telemetry vs. Survey Measurement for the family argument.
- Claude side — 1,909,132 global queries from the third AEI release (Appel et al., Aug 4–11 2025; 50.5% Free/Pro, 49.5% API), each mapped to one of 17,659 O*NET tasks and to one of five usage patterns. Tasks are sorted into ventiles by query volume and each ventile assigned a feasible-automation ceiling by a hand-set step function: ventile 20 → 80%, 19 → 70%, 18 → 60%, 17 → 50%, and the 85% of tasks with zero Claude queries → 5%, not 0 (the argument: generative AI can supply useful synthesis for "almost any task," including tree-trimming).
- The augmentation downgrade — the 223 tasks in the top 2% of the augmented-minus-automated query differential (all in the top two ventiles; mostly teaching, counseling, scholarly writing) are cut back from 70–80% to 40%, on the assumption that work users deliberately keep human-in-the-loop is intrinsically harder to hand over.
- ChatGPT side — OpenAI's 2025 workplace usage (Chatterji et al., May–Jul 2025) is public only across 41 Generalized Work Activities, not tasks, so a parallel decile→exposure schedule runs deliberately lower (top decile → 30%, floor 5%) to penalize the coarser aggregation, then weights each GWA by its O*NET importance to the job. Four GWAs OpenAI did not report are imputed from similar ones.
- Combine — a simple average of the two occupational scores.
The two halves behave nothing alike. Claude-only exposure runs 5%–62% (mean 14.5, SD 10.8) with a sharp left peak and strong positive skew; ChatGPT-only runs 27%–40% (mean 32.7, SD 2.2) and is bell-shaped. Same construct, same year, two frontier labs: one instrument spreads occupations across 57 points and the other across 13. Averaging them is the paper's way of splitting that difference.
The load-bearing step is the step function. Ventile→exposure and decile→exposure are assumptions the authors state and defend, not estimates. The measure inherits telemetry's ranking of tasks and takes its levels from a hand-set schedule; a different schedule moves every number in the column. Read it as "usage-ranked exposure with assumed ceilings," not as a measurement of feasibility. The construction also inverts the usual objection to usage data — the 85%-of-tasks-with-zero-queries floor is set by assumption precisely because absence of usage is not evidence of infeasibility.
Source-internal contradiction, and the caveat it leaves. §4.3 states the combined measure's range as "15.8 to 29.8" — while Table 2, the table that sentence cites, gives a maximum of 49.8. The table is the one to trust, on three independent grounds: the combined score averages a Claude component topping out at 62% with a ChatGPT component topping out at 40%, so its ceiling must sit near 51 and cannot be 29.8; the same arithmetic reproduces the floor exactly ((5+27)/2 = 16 ≈ 15.8); and the value was verified against the PDF at ingest, so this is the paper contradicting itself rather than a parse artifact (most likely a 4→2 transposition). The digit is not the point. The point is that this paper's prose summary statistics cannot be quoted without checking them against its tables, and this one cleared arXiv. Every figure on this page comes from Tables 1–4 or from prose that reconciles with them.
Averaging instruments, and what falls out#
The paper's answer to the disagreement is diversification rather than adjudication: average the standardized scores across models. Two exclusions along the way, both judgment calls, both disclosed:
- Massenkoff & McCrory dropped from every downstream analysis, because ρ=0.89 with Steele & Cruz would double-weight 2025 Anthropic usage. Six instruments remain in Figures 7–10 and the most-exposed table above.
- Frey & Osborne dropped from the headline average as "the oldest and most anomalous." Five instruments remain — though the paper reports that keeping all six "yields quite similar exposure recommendations overall," and shows the robustness check.
Worth stating plainly even so: dropping the instrument that disagrees most is the move that makes an average look stable. The reported convergence is partly a consequence of which instruments were kept. And a gap the paper never addresses: coverage runs 689–872 occupations, and the text nowhere states how the cross-model average handles occupations missing from some instruments — for a job absent from Frey, Webb, and Felten, a "five-model average" is a three-model average, unremarked.
What the average produces as career guidance (the paper's actual purpose, with an interactive tool):
- Median salary rises only modestly with cross-model exposure, and the highest-paying jobs (physicians, chief executives) sit near the exposure median, not at either extreme. Below-median salaries compress into a narrow $28K–$58K band while above-median pay spreads widely.
- Above-median pay with below-median exposure — the "stability" quadrant — concentrates in healthcare practice (the strongest single field), O*NET Job Zone 3 (associate-degree and skilled trades: plumbers, carpenters, dental hygienists, phlebotomists), and the Realistic, Social, and Investigative RIASEC categories.
- Above-median pay with above-median exposure covers management, finance, computing, engineering, law, and education — "fields that have been thought of as relatively reliable pathways in recent decades."
- Cross-model exposure peaks at the bachelor's-degree level, not above it.
Social jobs, and an assumption two older instruments got wrong#
The query-built instrument produces one result the authors themselves call counterintuitive: Social jobs (teaching, coaching, advising) score the highest exposure of any RIASEC category, because lesson planning and concept explanation are among the most common Claude queries. That cuts directly against an assumption built into Frey & Osborne 2017 and Brynjolfsson 2018 — that interpersonal work is an intrinsic human advantage and therefore an automation bottleneck.
The resolution is the augmentation/automation split rather than a flat refutation. Consulting, advising, and teaching are among the few GWAs where augmented queries outnumber automated ones — which is exactly why the paper's own instrument downgrades those tasks to 40%. The older instruments were right that this work resists full handover and wrong that it therefore resists AI at all. Their bottleneck intuition survives elsewhere intact: directing others, caring for others, negotiating, relating, developing teams, and coordinating work are rarely sent to either model.
And now one of those residual bottlenecks has been tested directly. Every instrument on this page infers exposure — from query counts, patents, raters, or rubrics. Jabarian & Henkel (arXiv 2607.28222) instead randomized 67,056 job applicants between an AI voice agent and human recruiters conducting the interview, and measured firm outcomes. Job interviewing is squarely the kind of relational, conversational, judgment-laden task Frey & Osborne and Brynjolfsson encode as an automation bottleneck, and which usage data alone cannot adjudicate — a task can go unqueried because it is infeasible or because nobody has built the interface. The AI arm produced 12% more job offers, ~18% more job starts and one-month retention, and no productivity decline, with human recruiters still making every hiring decision.
Three qualifications keep this from being a refutation of the bottleneck idea. The task automated was information collection, not the judgment — evaluation stayed human by design, so the experiment splits a "social" occupation rather than replacing it. The paper's own boundary condition preserves the older instruments' intuition where it plausibly holds: gains concentrate in "high-volume, high-turnover environments where tasks are repetitive [and] outcomes are rapidly observable," while "tacit knowledge, relational inference, or screening of highly specialized skills may benefit more from human screening." And a quarter of the human-arm interviews were screen-outs the AI did not perform, so part of the effect may be removed discretion rather than better information. Still, the direction matters for this page: the one relational task in the vault with a randomized outcome measure came out against the bottleneck assumption, and it did so on a firm's revenue metric rather than a query count.
Augmentation pays#
Ranking 872 occupations into deciles by share of augmented vs. automated Claude queries and plotting mean occupational salary: at the ninth decile, jobs using Claude heavily as a helper are about $7,000/yr better paid than jobs using it heavily as a delegate (the gap narrows at the tenth). Salary rises with Claude-usage decile throughout.
The authors under-sell it deliberately and the two caveats are real: the September 2025 release cannot separate work use from personal use, so tasks are simply assigned to the jobs that draw on them; and the salary data are from 2022, predating any AI-driven compensation change. A third, which they flag as timing: the data are post-Claude-Code-launch (May 2025) but predate its early-2026 autonomy improvements. Directionally it lines up with Returns to Expertise in Agentic Coding and Market-Priced AI Exposure (the AI Premium) — pay tracks the work that is hard to hand over — and the paper's closing hedge is the honest one: whether the pattern holds "will depend on usage norms adopted in each field."
Elite consensus, sitting on top of instruments that cannot agree on the sign#
On July 13, 2026 — three days before Steele & Cruz posted the head-to-head above — more than 200 economists and AI researchers, sixteen of them Nobel laureates, signed "We Must Act Now: A Statement on AI's Transformation of the Economy", warning of an economic transformation larger than the Industrial Revolution on a vastly shorter timeline. It is practitioner-opinion and carries no measurement; the vault holds it as a dated marker of elite opinion, not as evidence about anything on this page.
Put next to this page it is a useful marker precisely because of the mismatch in resolution. At the moment the profession's most senior figures signed a joint statement about AI's economic transformation, the seven instruments built to measure that transformation could not agree on which occupations are most exposed (twelve most-exposed slots, eleven distinct jobs), could not agree on whether exposure rises or falls with salary (positive for the four newest, negative for Frey 2017, flat-to-negative for Brynjolfsson 2018), and — holding a single instrument fixed — could not hold the sign of the exposure-to-outcome relationship stable across two overlapping time windows.
The overlap is not incidental. Brynjolfsson organized the letter and authored instrument #6, whose SML rubric produces one of the two anomalous salary gradients on this page and encodes the interpersonal-bottleneck assumption Controlled Variance: AI's Edge as Reduced Dispersion tested and found wanting. That is not a mark against the letter — it is the shape of the situation. Consensus on direction and urgency is compatible with total disagreement about magnitude and incidence, and the disagreement lives inside the signatory list, not between it and some outside critic.
The letter says as much where it is most careful, and the careful lines are the citable ones: Spence on "a high level of uncertainty about the magnitude and timing of the impacts across many parts of the economy," and Cunningham's "we are driving in the fog, and it is extraordinarily difficult to anticipate what will happen next." This page is what the fog looks like measured. The methodological reading: an appeal to expert consensus settles nothing this page is about, because the experts' own instruments are the thing in dispute — and the number of signatories is orthogonal to the resolution of the measurements.
Connections#
- Controlled Variance: AI's Edge as Reduced Dispersion — the outcome-measured counterpart to every instrument on this page: instead of inferring how exposed a relational task is, it randomizes 67,056 applicants through an AI-conducted job interview and reads offers, job starts, and retention off firm records. Direct evidence against the interpersonal-bottleneck assumption in Frey & Osborne 2017 and Brynjolfsson 2018, bounded by the fact that only information collection was automated
- The Solo-Authorship Rebound — an exposure ordering built from an outcome instead of a task rating, over a different unit: 26 scientific fields ranked by how much their solo-authorship trend broke after ChatGPT, which the author reads as the substitutability of a coauthor's execution work (Engineering +2.5 pp/yr, Business +1.9, down to Physics −0.0 and Arts and Humanities −0.5). It also restates this page's problem from inside another domain — Matsui concedes the occupational instruments (Eloundou 2024, Felten 2021) "do not map cleanly onto the OpenAlex fields," so a behavioral exposure measure and the task-rated ones cannot be joined even where both exist
- The Tragedy of the Cognitive Commons — exposure on a different axis: not what share of an occupation's tasks AI can do, but whether the tasks it takes are the ones novices learn on. Its five vulnerability factors predict where that matters most
- Task Crossover — the assumption under all four measures, challenged: exposure is computed per occupation against a fixed O*NET task list, but 43.5% of occupation-specific AI use is another occupation's work, so the tasks a job "has" are being reallocated faster than the mapping updates
- Anthropic Economic Index — the research program; observed/theoretical exposure are its earlier primitives, reported/anticipated the survey additions
- Task Saturation: Broad but Shallow AI Diffusion — a depth measure orthogonal to all four exposure measures here: not what share of tasks AI touches across the economy, but what share of a reached occupation's tasks it covers (median 21%) — and a second lab's observed exposure that lands at half the AEI's level
- Usage-Telemetry Classifier Validation — the classifier error bar under every exposure number, quantified for the first time by ATLAS: 22.6% exact accuracy at the O*NET task level against 85.8% human approval
- Returns to Expertise in Agentic Coding — the experience gradient is this survey's echo of the returns-to-expertise finding: tacit/relational judgment is the residual humans hold
- Organizational Complements to AI — the GDP gradient and the IMF complements argument: exposure ≠ impact without the complementary skills and infrastructure. Also the home of the HAT substitution model, the theory-side counterpart to this page's instrument disagreement: every instrument here measures what AI could do to an occupation, while HAT models the organizational decision to replace — and its own skill threshold
μ*turns out to be as sign-indeterminate as the instruments are - The Automation–Optimism Link — reported/anticipated exposure rise with automation share; the sentiment companion to this capability-belief measure
- Conversation Artifacts — artifacts are the output-side view of "what AI does"; complements the task-side exposure measures
- Conversation-to-Delegation Shift — the intensive-margin usage evidence behind observed exposure rising
- AI Usage Cadences — the off-hours/high-wage and cross-country usage rhythms foreshadow the GDP and occupation gradients this page measures directly
- Task Time-Horizon Scaling — theoretical exposure's ceiling is bounded by the reliable-task-length frontier
- Experimental Learning Impact of Generative AI — the belief-calibration analogue: like reported/anticipated exposure, students' perceptions of AI's learning effect are directionally right but miscalibrated in magnitude (control students overestimate the gain ~5×) until firsthand use corrects them
- Firm AI-Spend Intensity and Headcount Growth — also the home of the job-postings instrument whose window-dependent sign flip is described above, and of its contradiction with the Ramp panel on seniority composition. Beyond that, a measure on a different unit entirely: all four measures here are occupation-level exposure (what AI could/does do to a job); Ramp × Revelio observes firm-level adoption (which firms actually paid AI vendors, when, how much) linked to those firms' workforce outcomes — the firm-vs-occupation variation the paper argues exposure indices structurally cannot capture (two firms with identical workers can differ sharply in adoption). Steele & Cruz add the companion limitation: those indices cannot separate each other either
- Market-Priced AI Exposure (the AI Premium) — a fifth measure on a different axis: market-implied exposure, inferred from equity-price comovement with 380T tokens of realized AI consumption. Neither survey (reported/anticipated) nor usage/task-mapping (observed/theoretical) — it prices what investors believe about a firm's AI winners/losers, and its skill map (interaction/communication +0.36 SD, science most negative) is orthogonal to every task-based exposure measure (<2% of variance). It corroborates this page's experience gradient from the market side (relational/interactive skills priced up), but inherits its own developer skew rather than resolving the representativeness problem. Steele & Cruz's head-to-head reframes that orthogonality: the task-based measures barely agree with each other either, so being uncorrelated with them is not a mark against the market measure
- AI and Market Power — a new aggregation axis rather than a new construct, and the one this page said exposure indices structurally could not produce: OECD take the Felten-Raj-Seamans occupational GenAI score and employment-weight it up to the firm using Portuguese linked employer-employee data, generating firm-level variation from an occupation-level instrument. What comes out is two gradients pointing opposite ways — monotone in productivity, markup and tertiary share, inverted-U in size and market share. The reflexivity caveat is the sharpest version of this page's own: a score loading on cognitive non-routine tasks will find high-skill firms by construction, so the education gradient is partly the instrument looking at itself, while the size/market-share shape is not
- Telemetry vs. Survey Measurement — the family argument behind the seven-instrument clustering: Steele & Cruz's seventh model is the first to run usage telemetry forward into a projection (ventile-ranked tasks × assumed automation ceilings) rather than reporting it as an observed-exposure count, and the two telemetry-built instruments correlate at ρ=0.89 with each other while correlating poorly with everything task-based — measurement family, not construct, is what predicts agreement
Open Questions#
- Binned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the report checks robustness with a ≥60%-of-tasks indicator, but the levels remain self-reported)?
- Reported exposure exceeds observed partly because the survey reaches heavy users; what does the reported/observed gap look like in a representative sample? Sharpened: realized-consumption measurement (Borri-Liu-Tsyvinski) adds a market-implied instrument built on 380T tokens of actual paid requests — but it is skewed toward developers/sophisticated users (OpenRouter is ~2% of global tokens), a different non-representativeness than the survey's heavy-user skew. The lesson: no current AI-exposure instrument is representative; each collection mechanism biases in its own direction, so the reported/observed/market-implied gaps are partly artifacts of who each method reaches. A representative census remains the open target. Sharpened again by Steele & Cruz, which makes the population question concrete across seven instruments at once — 2,000 MTurk respondents (Felten), GPT-4 as rater (Eloundou), whoever files AI patents (Webb), Crowdflower workers against a rubric (Brynjolfsson), 70 jobs hand-coded by two researchers (Frey), and Claude/ChatGPT users in 2025 (Massenkoff; Steele & Cruz). Every instrument biases toward whoever it reaches, and the head-to-head shows the consequence is not a level shift that a rescaling would fix: the instruments produce different job rankings, nearly disjoint most-exposed lists, and opposite signs on the exposure-salary gradient.
- Does averaging across instruments reduce error or merely blend incompatible biases? Steele & Cruz's cross-model average is a diversification argument, not a validated one — no instrument in the set has been scored against realized labor-market outcomes, and the average's apparent stability partly reflects dropping the two instruments that disagreed most (Massenkoff for redundancy, Frey for anomaly). Falsifiable: score all seven, plus the average, against subsequent occupation-level employment and wage changes. Complication (2026-08-04): Indeed's postings data shows the exposure–outcome relationship changing sign between windows on a single instrument (2022–2026 negative, 2025–2026 positive), so any such validation scores the window as much as the instrument, and a scoring period must be pre-specified rather than chosen after the fact.
- The experience gradient rests on what workers believe AI can't do (judgment, relational work) — a belief that could be either durable comparative advantage or the next capability to fall. Which, and when?
Sources#
- Anthropic Economic Index report: Cadences — Anthropic Economic Index report: Cadences (June 26, 2026), Chapter 3 "Perceptions": §AI and work tasks (Figures 3.2–3.4)
- Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Brian Jabarian & Luca Henkel, Voice AI in Firms (arXiv 2607.28222, 2026-07-30;
empirical, pre-registered RCT). §3.1 (offer/start/retention effects), §3.2 (productivity nulls), §4.1 (the 25%-vs-7% screen-out gap), §8 (boundary conditions on where AI screening helps). Cited here as the outcome-measured test of the interpersonal-bottleneck assumption; parse warnings and full treatment at Controlled Variance: AI's Edge as Reduced Dispersion. - The Human-AI Substitution Principle: When will you be replaced by AI in your organization? — Banerjee & Singh, arXiv 2607.20781 (2026-07-22;
practitioner-opinion, formal model, no empirical data): §4.3.2 Theorem 9 + Corollary 3 (theμ*protection/vulnerability threshold and its explicit refusal to sign the skill relation without domain-specific parameters), §5.6 Table 3 prediction P5 and §5.6.5 (the "counterintuitively, high-skill workers may be more vulnerable" claim), §3.4 Assumption 1(i)Δ$' ≤ Δ$. Full treatment at Organizational Complements to AI - AI and Job Postings: From Destruction to Creation? — Guillermo Gallacher, AI and Job Postings: From Destruction to Creation? (Indeed Hiring Lab, 2026-07-08;
empiricalpostings data). §"Are other occupations exposed to AI experiencing a similar rebound?" — the May 2022 → May 2026 negative and May 2025 → May 2026 positive exposure–postings relationships, the exposure measure sourced to Hiring Lab's own AI at Work Report 2025, and the cited New York Fed finding that the AI-exposed vacancy decline predates ChatGPT. COI: Indeed's research arm analyzing Indeed's own job board, published as a blog post. Chart-only, therefore not quoted anywhere in this vault: both scatter plots' correlation coefficients and p-values (the prose asserts statistical significance without reporting either), and the six countries and per-country shares in the international chart (prose names only Germany and France as exceptions). Full evidence note at Firm AI-Spend Intensity and Headcount Growth. - “We Must Act Now”: Sixteen Nobel Laureates Join Leading Economists and AI Researchers in Call to Prepare for AI’s Economic Transformation — Matty Smith, Stanford Digital Economy Lab news release, 2026-07-13 (
practitioner-opinion, 558 words, no measurement): the 200+ signatories / sixteen Nobel laureates count, the four organizers, and the Spence and Cunningham uncertainty quotes. Aggregate counts only — the statement text and the live signatory list live atwemustactnow.ai, a separate uningested source that has since grown well past the count reported here. Cited on this page as a dated marker of elite opinion, never as evidence about exposure - Helping People Choose Careers in the Age of AI — Jennifer L. Steele & Isabella Cruz, Helping People Choose Careers in the Age of AI (arXiv 2607.15506, 2026-07-16; American University / CU Boulder;
empirical). §3 model descriptions + Table 1 (the seven instruments); Table 2 (raw distributions); §4 + Tables 3–4 (the query-based construction); §4.4 + Fig. 6 (correspondence); §5.1–5.2 + Figs 7–8 (salary and complexity gradients); Table 5 (most-exposed occupations); §5.3 + Fig. 10 (RIASEC); §5.5 + Figs 13–16 (salary × exposure quadrants); §5.6 + Fig. 17 (augmented vs. automated salary). Parse warnings. Tables 1–4 verified clean cell-for-cell at ingest and are quoted directly. Table 5 is collapsed — all six model names merged into row 1 with rows 2–6 blank, occupation columns intact; the mapping reproduced above was recovered from the PDF, never from the parsed table. Appendix Table A1 is collapsed and row-shifted (column-3 coefficients slid down one row) and is not cited on this page. Source-internal contradiction: §4.3's prose range "15.8 to 29.8" contradicts Table 2's max of 49.8 for the same measure; the table is correct on arithmetic and PDF verification — treat this paper's prose summary statistics as needing a table check.
Cited by 23
- Anthropic Economic Index×5
The AEI publishes its task-level data, and by mid-2026 outside economists were building independent…
- Market-Priced AI Exposure (the AI Premium)×5
Its methodological significance for this vault: it is a third measurement paradigm for AI exposure.…
- Task Crossover×4
The most consequential claim, and it applies to nearly every number in Exposure Taxonomy. Observed,…
- Firm AI-Spend Intensity and Headcount Growth×3
Exposure Taxonomy — the axis this paper adds: the taxonomy's four measures are all occupation-level…
- Telemetry vs. Survey Measurement×3
Two consequences for this page. First, the telemetry/survey split is not just a latency difference…
- AI and Market Power×2
Because no official statistics capture GenAI use yet, the paper proxies potential GenAI use with a…
- Erik Brynjolfsson×2
Exposure Taxonomy — author of instrument #6 (SML 2018), one of the two whose exposure–salary…
- Open Questions Backlog×2
Exposure Taxonomy ×3 (oldest 41d) — Binned midpoint coding biases the exposure slopes toward zero;…
- Organizational Complements to AI×2
Exposure Taxonomy — the GDP gradient in reported AI exposure is this page's argument in survey…
- Returns to Expertise in Agentic Coding×2
Exposure Taxonomy — the AEI Cadences survey confirms this from the worker's mouth: 15+-year workers…
- The Solo-Authorship Rebound×2
It is a field-level exposure ordering built from behavior, which is what Exposure Taxonomy finds…
- AI Usage Cadences
Exposure Taxonomy — cross-country and off-hours patterns foreshadow the GDP/occupation gradients…
- The Automation–Optimism Link
Exposure Taxonomy — reported and anticipated exposure also rise with automation share; this page is…
- Controlled Variance: AI's Edge as Reduced Dispersion
Exposure Taxonomy — a direct experimental test of the interpersonal-bottleneck assumption baked…
- Conversation Artifacts
Exposure Taxonomy — artifacts are the output side of what "AI can do a task" means; complements the…
- Conversation-to-Delegation Shift
Exposure Taxonomy — the survey's belief-side measures, complementing this behavior-side shift
- Experimental Learning Impact of Generative AI
Exposure Taxonomy — the belief half: like reported/anticipated exposure, students' perceptions of…
- Google AI & Economy ATLAS
Exposure Taxonomy — ATLAS supplies observed exposure from a second lab; its task-saturation measure…
- AI Economics & Labor
Exposure Taxonomy — Four distinct ways to measure AI's reach into an occupation — observed exposure…
- Task Saturation: Broad but Shallow AI Diffusion
Exposure Taxonomy — task saturation is a stricter observed exposure measure from a second lab; the…
- Task Time-Horizon Scaling
Exposure Taxonomy — theoretical exposure (what an LLM could do) is bounded above by this…
- The Tragedy of the Cognitive Commons
Exposure Taxonomy — the five vulnerability factors are an exposure measure on a different axis: not…
- Usage-Telemetry Classifier Validation
Exposure Taxonomy — every exposure measure is a classifier output; this quantifies the error bar…
Related articles
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Organizational Complements to AI
The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…
- Conversation-to-Delegation Shift
OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production…
- The Automation–Optimism Link
AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-qualit…
- Telemetry vs. Survey Measurement
Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (…
