H
Howardism
Plate IIAI Economics & LaborHOWARDISM

Does the Augmentation/Automation Split Govern Skill at Work?

Verdict: the split transfers to the workplace in sign but is unestablished as a measurement. Four workplace deskilling results sit on the automation row (Budzyń's endoscopists, Dell'Acqua's consultants, Vicente & Matute's 80.7%, Wiles et al.) and two on the augmentation row (Everett's collaborative clinical workflows, Brynjolfsson's support agents) — but all six reach the vault secondhand through one `practitioner-opinion` framework paper, none classifies use mode within a design, and none takes an unaided post-measure after AI is withdrawn. The disanalogy that bites is not the elite sample or the one-week horizon: the lab's mechanism is reallocation with time-on-task *fixed* (−5.3pp writing / +4.4pp reading), while at work AI's headline effect is time saved and nothing in the vault measures what the freed hour buys. Workplace volume also produces a third mode a 35-minute proctored session cannot generate — skipped review (31.3% of PRs merged unreviewed; 81.1% of genuine leaked credentials drawing no reviewer comment) — which is not the automation arm, since the automating students still submitted work they had read. The AEI cannot arbitrate because 'automation share' welds Directive to Feedback Loop and excludes the classifier's own Learning mode by construction, and that axis is drifting toward the hollow arm (43–45% and rising, against the lab's 32%). The one randomized workplace case runs the other way: controlled-variance's recruiters redeployed expertise onto the un-automated stage (discounting the AI's interview signal, −0.047/−0.029) and paid in latency (20→24 days), not skill.

Article metadata
Publication details
Published:September 5, 2026
Filed:Essay
Domain:AI Economics & Labor
Reading:22 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Does the Augmentation/Automation Split Govern Skill at Work?

The question#

The #oq/now bullet on Experimental Learning Impact of Generative AI:

Does the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer?

The split to be transferred is the load-bearing result of Contractor & Reyes's randomized experiment: classifying treated students from their ChatGPT logs, augmentation users (AI as tutor/explainer, 49%) hold +0.29 SD on an unaided test a week later, while automation users (AI drafts the text, 32%) post +0.54 SD on the session-one essay and +0.02 SD on the unaided session-two essay — the short-run gain is AI-produced output that does not survive the tool's removal.

The answer in short#

  1. The sign transfers. The vault holds four workplace or professional results that sit squarely on the automation row and two on the augmentation row, from settings the lab does not resemble. The pattern replicates outside a proctored task.
  2. The measurement does not. Every one of those six reaches this vault secondhand, through The Tragedy of the Cognitive Commons — a practitioner-opinion framework paper whose author labels his own mechanism "a structural prediction … not an observed outcome." None classifies use mode within a design; none takes an unaided post-measure after AI is withdrawn. What transfers is a pattern assembled across studies, not a replicated coefficient.
  3. The disanalogy that bites is time, not population. The experiment's stated mechanism is reallocation with total learning time held fixed−5.3pp writing / +4.4pp reading-and-searching, and explicitly no effect on time spent. At work, time saved is AI's headline effect, and the authors say plainly that this is why they do not claim real-world learning gains. Nothing in the vault measures what the freed hour buys.
  4. The workplace has a third mode the lab cannot produce. AI Brain Fry (+11% minor / +39% major errors), Acceleration Whiplash (31.3% of PRs merged with no review), and Security Debt of Agent-Generated Code (81.1% of genuine leaked credentials reaching integration uncommented) describe workers who skip the review step entirely under volume. That is not the automation arm — the automating students still submitted an essay they had read — and a 35-minute proctored session cannot generate it.
  5. The AEI cannot arbitrate, for a reason the experiment identifies. The Automation–Optimism Link's "automation share" is Directive + Feedback Loop, which pools the two types that diverge, and the same classifier's Learning mode is excluded from the axis by construction. A flat self-reported learning curve is therefore compatible with a real hollow half.
  6. The mix is drifting the wrong way. The AEI's automation share is 43–45% and rising (as ATLAS notes when disputing the level), against the lab's 32%. If the split governs, the workplace mix is moving toward the hollow arm.
  7. The one randomized workplace case runs the other way. In Controlled Variance: AI's Edge as Reduced Dispersion, recruiters whose interviewing task was automated did not hollow out — they redeployed judgment onto the stage that remained, discounting the AI's interview signal and leaning on independent tests. The cost was latency, not skill.

Verdict: partially answered. The split is the best available hypothesis for workplace skill accumulation and has converging support in sign, but it is not established at work by anything in this vault, and two of the transfer's premises (fixed time-on-task, no volume pressure) fail by construction outside the lab.

1. What transfers: the sign, assembled across four workplace results#

The commons framework's evidence ladder (The Tragedy of the Cognitive Commons) is where the vault's workplace-side skill evidence lives. Sorted onto the experiment's two rows, it lines up:

Study (all secondhand via the commons paper)SettingRow
Budzyń et al. (2025), Lancet Gastro & Hepatology — endoscopists' independent detection accuracy fell after adopting AI-assisted detectionScreening colonoscopy, career-length practiceAutomation
Vicente & Matute (2023) — 80.7% of participants detected the errors in biased AI recommendations, followed them anyway, then reproduced the bias in their own later judgments after the AI was removedLab, but the removal design is the experiment's ownAutomation
Wiles et al. (2024) — AI assistance raised performance during access and produced no significant advantage on later unassisted assessmentAssisted task with unaided post-testAutomation
Dell'Acqua et al. (2026) — elite consultants gained inside the AI capability frontier and lost outside it, unable to tell which side a task sat onProfessional consultingAutomation
Everett et al. (2025) — collaborative AI workflows requiring active clinical engagement improved diagnostic accuracyClinicalAugmentation
Brynjolfsson's support agents learn (cited on Experimental Learning Impact of Generative AI as the augmentation-side workplace analogue)Customer supportAugmentation

Two things this table earns. First, the removal signature — performance high with the tool, no advantage without it — is exactly the automation row's shape, and Budzyń, Wiles, and Vicente & Matute each independently produce it on non-students. Second, the split's direction is not an artifact of the classroom: Everett's result is the augmentation row in a safety-critical professional setting, and it is the difference in whether the practitioner's own attempt is preserved that separates it from Budzyń's, which is the same variable the ChatGPT-log classifier reads.

The commons paper reaches the same conclusion from theory and states the conditionality in the experiment's own terms: AI reallocates cognitive effort from generative struggle toward orchestration and evaluation, and "whether that reallocation builds or erodes expertise is conditional on how the work is designed." That is the use-mode split, stated as an organizational-design variable rather than a per-conversation one.

2. What does not transfer: nobody splits use mode inside a workplace design#

The experiment's power is not the effect size. It is that one randomized population, on one outcome instrument, was partitioned by its own logs into two use modes whose measured skill trajectories diverge — validated internally (automation users' essays 53.6% AI-flagged vs 20.9%, and 17% less time reading/searching). Nothing in the vault does that at work:

  • Budzyń, Dell'Acqua, Wiles, Vicente & Matute, and Everett are five separate studies of five separate settings, and they arrive here entirely through a conceptual paper. The vault has read none of them directly; every figure is the commons paper's rendering. This is not enough to license a claim about a coefficient.
  • The commons paper's own boundary conditions cut against over-applying the assembly: it acknowledges null effects on earnings and hours in Denmark across two years of adoption, adaptive capacity in aggregate labor data (Manning & Aguirre 2026), and heterogeneous retrainability (Hyman et al. 2025) — "these findings establish that the depletion mechanism is not universal."
  • The vault's own strongest workplace panels measure the wrong quantity. Returns to Expertise in Agentic Coding measures what a stock of expertise buys you from an agent (+9% actions / +13% output per expertise level; verified success 15% novice → 28–33% intermediate-expert, a concave curve where most of the gain is novice→intermediate). That is expertise as an input, not accumulation as an output — it tells you what a deskilled worker would lose, not whether AI use is deskilling them.

3. The disanalogy that actually bites is time, not population#

The obvious objections — elite liberal-arts undergraduates (GPA 3.68, SAT 1386), 35 minutes, one week, one topic — are real but not the ones that break the transfer. The one that does is structural, and the paper names it:

It measures learning per unit of time with time-on-task held fixed. AI changed how the hour was spent (−5.3pp writing / +4.4pp reading-and-searching), not its length — plausibly because the lab offered few competing uses of time. The authors explicitly do not claim that real-world AI adoption raises learning overall: outside the lab students choose how long to study, and many use AI to save time.

At work the fixed quantity is output, not time, and the vault's workplace instruments measure exactly the time savings the lab suppressed — Noy & Zhang's writing gains, the Acceleration Whiplash throughput rise (tasks +34%, epics +66%). The mechanism the experiment identifies (effort moves from producing text to absorbing it) requires that the hour stays put. Where the hour is spent elsewhere, the augmentation dividend is not measured by this design in either direction. This is the first of Experimental Learning Impact of Generative AI's own open questions, still #oq/source, and the workplace question inherits it.

The one workplace decomposition in the vault that speaks to this runs against the intuition, which is worth recording: in Controlled Variance: AI's Edge as Reduced Dispersion, automating the interview did not free the recruiter's time, it consumed more of it — the interview→offer-decision stage went from a median 2.62 to 7.24 days (p=0.005) because recruiters now review a conversation they did not have. End-to-end time-to-hire rose 20 → 24 days. Where the human retains the evaluation stage, delegation can buy more engaged time with the residual judgment task, not less.

4. The workplace has a third mode the lab cannot generate: skipped review#

This is the sharpest reason the transfer is incomplete. The experiment's automation arm is a student who let AI draft an essay they still read, edited, and submitted under proctoring, in a session with no volume pressure. The vault's workplace measurements of AI-era cognitive load describe something else:

  • AI Brain Fry — Kropp et al.'s mental fatigue from oversight beyond cognitive capacity: workers reporting brain fry make 11% more minor and 39% more major errors. The page's own framing is that the tool framing taxes the reviewer while the employee framing replaces the tax with under-engagement — a distinct failure mode with the 18% drop in error-catching attached.
  • Acceleration Whiplash — Faros's org-scale telemetry across ~4,000 teams: daily PR contexts per developer +67.4%, work restarts +13.8%, and 31.3% of PRs merged with no review at all.
  • Security Debt of Agent-Generated Code — inside agent-authored PRs, humans committed 67.6% of the 74 genuine leaked credentials and 81.1% of them reached integration with no comment from any bot or human reviewer, which the authors read as reduced vigilance in workflows where the agent appears to be handling correctness.

Call this abdication, and note that it is not the automation row. Automation-mode use in the experiment still routes the output through the human, which is why its short-run gain is real (+0.54 SD) even though it is hollow. Abdication removes the human from the loop, so it produces neither the short-run gain nor the learning — and it is generated by volume, the variable a 35-minute session holds at one. If the workplace question is "does AI use erode skill," the lab supplies one mechanism (hollow output-substitution) and the workplace telemetry supplies a second the lab cannot see. Both are consistent with Outsource Your Thinking, Not Your Understanding; only the first is what the experiment measured.

5. Why the AEI survey cannot arbitrate — its own axis pools the split#

The Automation–Optimism Link is the workplace instrument that should have answered this, and it reports a clean null: heavy delegators report learning at the same rate as everyone else, 68% report learning more, 57% report AI makes their skills more valuable, and optimism rises with automation share across all six job-quality dimensions. Three reasons this is not evidence about accumulation:

  1. The axis welds the two arms. Automation share is the fraction of conversations classified Directive or Feedback Loop. Directive is the experiment's automation arm almost exactly ("human delegates the whole task, minimal interaction"). Feedback Loop — "iterative, human mainly feeds back from the environment" — is not obviously either. And the classifier's own Learning mode ("human seeks understanding, not task completion") is the experiment's augmentation arm, and it is excluded from automation share by construction. So the survey's headline is computed on an axis that already contains the distinction and then collapses it. A flat learning curve across that axis is exactly what a real hollow half would look like when averaged with a real solid half.
  2. Self-report is the wrong instrument, and the lab shows it is miscalibrated in both directions. In the experiment, self-assessed knowledge showed no treatment effect (β=0.02) while measured knowledge rose +0.27 SD — students learned more and did not feel it. Control students predicted a +25.2pp own gain against an actual 5.1pp, a 5× overestimate that firsthand use corrected to +3.8pp. Felt and measured skill diverge in magnitude and sign depending on whether the respondent has used the tool. The AEI page states the caveat itself: skills can erode even as people feel they are learning.
  3. The sample and the identification. ~9,700 self-selected Claude users, computer/math ~30%, management ~23%, women 12%; correlational, with the report explicit that selection is not fully ruled out.

This closes a loop that was already half-closed: The Automation–Optimism Link's own second open question ("is there an objective skill measure that agrees?") was marked Partially answered by the experiment. What this synthesis adds is that the pooling is not incidental — it is a property of the AEI's collaboration-mode classifier, so no re-analysis of the existing survey can separate the arms; the automation-share axis would have to be decomposed into Directive vs Feedback Loop, and Learning re-admitted as a use mode rather than a residual.

6. The mix, not the mechanism, is what is actually unknown at work — and it is drifting#

The experiment's third open question notes that the augmentation/automation choice is endogenous to incentives. Its own logs run 49% augmentation / 32% automation, with explaining concepts the single most common use (57%). Workplace telemetry runs the other way, and is moving:

  • Anthropic Economic Index: automation share 43–45%, and — the observation ATLAS uses when disputing the level — rising over time. ATLAS classifies <10% of non-routine-cognitive conversations as end-to-end automation intent, a gap the AEI page reads as mostly definitional (binary vs five-category, anything short of end-to-end counted as collaboration).
  • Conversation-to-Delegation Shift: the same asking→doing move in output tokens — Codex share 99.8% (OpenAI) / 63.3% (organizational) / 16.5% (individual), with the >8h-task share rising 2.1% → 25.6%.
  • Market-Priced AI Exposure (the AI Premium): cross-provider corroboration outside any single lab — agentic token share (finish reason tool_calls) from near zero in 2024 to 52.2% by early 2026.

If the split governs skill, the workplace mix is drifting toward the hollow arm while the classroom mix is not. Two caveats keep this directional rather than decisive: the two telemetry programs disagree on the level by a factor of four on definitional grounds, the AEI's classifier accuracy is unpublished (Usage-Telemetry Classifier Validation), and none of these instruments measures skill at all — they measure delegation. The drift is a reason to expect the question to matter more over time, not evidence about its answer.

7. The one randomized workplace case runs the other way#

Controlled Variance: AI's Edge as Reduced Dispersion is the vault's only randomized causal estimate of AI substituting for a human in an expert conversational task, and its recruiters are the closest thing the vault holds to a workplace population whose task was automated under experimental control. What happened to their judgment is not hollowing:

  • They scored AI-conducted interviews higher (mean 1.90 → 2.01, p<0.001, driven entirely by a shift from score 1 to score 2 — the AI lifts the floor, not the ceiling).
  • They weighted that signal less. Regressing offers on standardized interview, language-test and analytical-test scores, the interview score is significantly less predictive of offers in the AI arm (interaction −0.047, −0.029 with controls) while the language-test score is more predictive (+0.028 / +0.022) — and the discount is concentrated exactly among the recruiters who had told the survey that interview performance matters more than test scores.

Discounting a signal precisely where your own stated model says it should be discounted is expertise being redeployed onto the un-automated stage, and the paper reads it the same way: "realizing the full returns from automation requires complementary adaptation in how humans rely on AI signals for decision-making." That is Organizational Complements to AI's complement-that-didn't-get-built, and it is the augmentation-row outcome at the level of a job rather than a conversation.

Three things stop this from settling the question in the optimistic direction. No unaided skill measure was taken — nobody asked whether a recruiter who had not conducted an interview in three months could still conduct one, which is the exact measurement the transfer needs. The horizon is four months, and the four-month retention estimate is the one that fails recruiter clustering. And the recruiters' forecasts were wrong in sign — 36% expected AI-interviewed applicants to receive lower offer rates, 48% expected lower retention, 61% expected lower interview quality — so incumbent expertise mispredicted the outcome even as it correctly reweighted the signal.

One further datum belongs to the workplace half of the automation–optimism link and inverts it: only 12% of recruiters expected AI's impact on themselves to be generally positive, against 47% of applicants, inside the same firm. The AEI's finding that delegation breeds optimism is measured on people delegating; the person whose task is the one being automated is the case the survey does not contain.

8. Verdict#

The use-mode split is the right hypothesis for the workplace and the vault cannot confirm it there. Precisely:

  • Transfers: the sign, on six studies of non-students, with the separating variable (is the practitioner's own attempt preserved?) matching the log classifier's.
  • Does not transfer: any within-design split, any unaided post-measure, any first-hand reading of the evidence — all six workplace results are secondhand through a practitioner-opinion paper.
  • Fails by construction: fixed time-on-task (the lab's whole mechanism) and single-task volume (which is what generates abdication, the workplace's own third mode).
  • Cannot be recovered from what exists: the AEI's automation share pools Directive with Feedback Loop and excludes Learning, so no re-analysis of the existing survey separates the arms.

So: not "too unlike to transfer," and not "settled." The question moves from #oq/now to #oq/source, because what it now needs is a workplace study that does not exist in the vault, not more synthesis over what does.

What would change this#

  • A workplace study with an unaided post-measure. Randomize or track AI access on a real task, classify use mode from logs within subject, then measure performance after withdrawal. This is the single missing design; everything else here is a proxy for it. The commons paper proposes the instrument independently — "no-AI performance assessment and error-detection audits" — without anyone having run it.
  • Decomposing the AEI's automation share. Directive vs Feedback Loop, with Learning re-admitted as a mode rather than a residual, against the linked survey's skill items. Anthropic already has both the classifier and the linked panel; this is a re-aggregation, not a new study. It would not fix the self-report problem, but it would say whether the flat learning curve survives un-pooling.
  • A within-person AEI panel (already The Automation–Optimism Link's own first open question): sentiment and a skill proxy tracked before and after adopting automated workflows would separate selection from treatment on the axis the drift in §6 is measured on.
  • A second round of Controlled Variance: AI's Edge as Reduced Dispersion's recruiter arm. If interview-score predictiveness recovers as recruiters adapt, the redeployment reading holds; if it degrades further while offers stay flat, the recruiters are rubber-stamping and the abdication mode has been caught under randomization.
  • A measurement of what the freed hour buys. The lab's mechanism dies the moment time is endogenous, and no vault source measures reallocation of AI-saved time at work. Until one does, §3 is the transfer's binding constraint.
  • First-hand ingest of Budzyń, Dell'Acqua, Wiles, or Everett. Four of the six load-bearing workplace results are the commons paper's rendering of other people's numbers. Reading any one directly would upgrade §1 from an assembled pattern to a citable measurement — and would test whether the row assignments here survive contact with the designs.

Evidence tiers and weighting#

  • Experimental Learning Impact of Generative AIExperimental Evidence on the Learning Impact of Generative AI (empirical, Contractor & Reyes, arXiv 2607.08849) — the anchor, and the only causal identification of the split anywhere in this synthesis: randomization, F-test balance, double-lasso controls, ITT + TOT/2SLS, proctored compliance, blind human + LLM grading. Discounts: 211 elite undergraduates, 35 minutes, one week, one model, one topic — and the augmentation/automation classification is post-hoc on logs, so the split itself is a heterogeneity analysis, not a randomized contrast. Its own honesty about time-on-task is what §3 is built on.
  • Controlled Variance: AI's Edge as Reduced DispersionVoice AI in Firms: A Natural Field Experiment on Automated Job Interviews (empirical, Jabarian & Henkel, arXiv 2607.28222) — the only randomized workplace evidence used here: pre-registered, 67,056 randomized applicants, administrative outcomes. Weighted highest for §7 and §3's latency decomposition. Discounts carried from the page: the AI system is never identified (no model, vendor, or version), the author's unpaid Chief Economist role at the partner firm postdates data collection but not the writing, the recruiter-side findings are observational within a randomized design, and there is no skill measure at all.
  • The Automation–Optimism LinkAnthropic Economic Index report: Cadences (empirical tier, but a self-reported survey on a non-representative sample, correlational) — weighted as evidence about perception, not accumulation, for the reasons in §5. The classifier underneath it has no published validation (Usage-Telemetry Classifier Validation), which is also the caveat on §6's drift.
  • The Tragedy of the Cognitive Commons (practitioner-opinion, conceptual, every figure secondhand) — the source of all six workplace studies in §1. Weighted as a bibliography with an argument, not as measurement: it is the vault's route to Budzyń, Dell'Acqua, Wiles, Vicente & Matute and Everett, and the framework's author states his central mechanism is "a structural prediction … not an observed outcome." Its acknowledged counter-evidence (Danish nulls, Manning & Aguirre, Hyman et al.) is weighted equally with its supporting cases, because it comes through the same hand.
  • AI Brain Fry → Kropp et al., HBR 2026/03, reaching the vault secondhand through the May 2026 HBR follow-up (Research: Why You Shouldn’t Treat AI Agents Like Employees, no evidence: tier — a pre-classification source). The +11%/+39% is workers reporting brain fry against error frequency: cross-sectional and correlational, not causal. Used in §4 for the mode it names, not for its magnitudes.
  • Acceleration Whiplash and Security Debt of Agent-Generated Code (both empirical, telemetry and static analysis respectively) — the strongest measured support for the abdication mode, and the strongest evidence in the synthesis that did not come through the commons paper. Neither measures skill; both measure the review step disappearing, which is why §4 treats abdication as a distinct mechanism rather than as evidence about the split.
  • Returns to Expertise in Agentic CodingAgentic coding and persistent returns to expertise (empirical, first-party Anthropic telemetry, ~400K sessions) — used only to establish that the vault's expertise measurements are of a stock, not of accumulation. First-party, transcript-inferred outcomes, excludes headless/SDK usage.

Tier ordering across the synthesis: the one randomized learning result and the one randomized workplace result carry the argument; the telemetry programs establish the drift but not the outcome; the survey establishes perception; the framework paper supplies the workplace pattern at the cost of being the weakest tier in the set, and §1 is stated with that discount attached.

§ end
Cited by 10
Related articles