Sources#
Summary#
Zara Contractor and Germán Reyes (Middlebury College, arXiv 2607.08849, July 2026) run the study most of the AI-and-learning corpus was missing: a randomized controlled experiment with an objective, proctored, unaided skill measure. 211 undergraduates learn an unfamiliar technical topic (blockchain, carbon capture, or CRISPR) and write an analytical essay in a 35-minute lab session, randomly assigned to AI-allowed or AI-forbidden; one week later all of them return and are tested and write again without any AI or resources. The headline: AI access raises immediate test scores by 0.27 SD, about 76% of that gain still shows up a week later unaided, and essay quality rises only after AI is removed. But the durable part is not uniform — it belongs almost entirely to students who used AI to explain concepts (augmentation), while students who used it to draft their text (automation) post large short-run essay gains that vanish completely once AI is gone. This is the near-controlled test of "outsource your thinking, not your understanding" and the objective-skill measure the automation–optimism link's self-report could not supply.
Evidence note.
empiricaland causal (randomization, F-test balance, double-lasso controls, ITT + TOT/2SLS) — the strongest design in this vault's learning cluster. But scope it honestly: elite liberal-arts undergraduates (mean GPA 3.68, SAT 1386, >80% already AI-adopters), a proctored 35-minute lab task, off-the-shelf ChatGPT (GPT-4o), one topic, one week. It measures learning per unit of time with time-on-task held fixed (AI changed how the hour was spent, not its length — plausibly because the lab offered few competing uses of time). The authors explicitly do not claim that real-world AI adoption raises learning overall: outside the lab students choose how long to study, and many use AI to save time. It informs, but does not settle, workplace skill-atrophy.
The design in one paragraph#
Between-subjects randomization at the lab level (eight time slots × two parallel labs). Session One: baseline 5-question test → 35-minute learning phase (read, search, draft a ~500-word essay; AI-allowed group gets a logged-in ChatGPT tab, AI-forbidden gets Google) → unaided 5-question post-test. Compliance enforced by proctors, ChatGPT logs, and interface screenshots. Session Two (~7 days later, everyone unaided): 10-question test + a 20-minute essay on the same topic, complementary prompt. Learning is measured two ways: knowledge tests (factual/conceptual recall) and essays (higher-order analysis), the latter graded blind by 311 master's/PhD Prolific graders plus an LLM grader, averaged, with objective linguistic features (length, readability, lexical diversity, textual similarity, and Pangram AI-detection) alongside.
First stage was strong: assignment lifted AI use by 67.3pp off a near-zero control; 68% of the treated used AI; 88% found it helpful. In the logs, the most common use was explaining concepts (57%), then drafting (31%), then summarizing (17%).
Durable test-score gains, concentrated in the middle — and skewed to the able#
- Immediate: +6.7pp on a 56.3% control baseline = +0.27 SD (ITT; TOT +10.0pp / +0.40 SD). More than double the 0.10 SD median of Kraft (2020)'s 747 education RCTs, comparable to the 0.29 SD of structured human tutoring, ≈ a one-SD jump in GPA, ≈ $8,600/student of school spending.
- Retention (one week, unaided): +5.1pp = +0.27 SD, ~76% of the immediate effect persists (the two are statistically indistinguishable). Durable, partially decaying knowledge.
- Shape: the gain sits in the middle of the score distribution — the tails (0/1 and near-perfect scorers) barely move (Figure 4).
- Who gains: larger in the upper GPA/SAT quartiles (bottom quartile ~0.05 SD; upper quartiles 0.19–0.40 SD) — so AI access may widen learning gaps, a genuine tension with the floor-raising story of software democratization.
- Felt vs. measured: self-assessed knowledge shows no treatment effect (β=0.02). Students learned more but did not feel they had — the mirror image of the felt-vs-measured worry the AEI self-report leaves open.
Higher-order skills surface only after AI is removed#
In Session One, treated essays are longer, simpler, and 12.3pp more AI-flagged (a 96% jump), but overall quality rises only slightly and imprecisely (+0.17 SD, p=0.25) — because the essay is partly AI-written, blending output with learning. Notably, no homogenization: within-group essay similarity is flat, contra the convergence Brynjolfsson et al. (2025) found in workplace writing (open-ended prompts admit many valid answers).
The clean read comes from Session Two, written unaided a week later (Figure 6): the AI-detection effect collapses to 0.0pp and the stylistic differences fade — proving the Session One style shift was AI text entering essays, not durable change — while quality gains emerge: writing style & clarity +0.30 SD (p=0.016) and relevance to prompt +0.26 SD (p=0.041). The learning was real and it raised higher-order skill, not just fact recall — but it only became visible once the AI crutch was taken away.
Augmentation vs. automation: the controlled test of "outsource thinking, not understanding"#
The paper's most load-bearing result. Classifying treated users from their ChatGPT logs: 49% augmentation (AI works with the student — tutor, explainer), 32% automation (AI does the work — drafts the essay), 8% mixed, 11% off-topic. Validation that the split is real: automation users' essays are 53.6% AI-flagged vs 20.9% for augmentation users, and automation users spend 17% less time reading/searching.
| Session 1 essay quality | Session 2 essay quality (unaided) | Session 2 test (unaided) | |
|---|---|---|---|
| Automation users | +0.54 SD (p=0.029) | +0.02 SD — vanishes | +0.19 SD (p=0.33, n.s.) |
| Augmentation users | +0.05 SD | +0.22 SD (p=0.22) | +0.29 SD (p=0.06) |
Automation's Session One advantage is AI-produced output that does not survive its removal; augmentation's gains persist unaided, consistent with skill accumulation. This is a within-experiment demonstration that AI's short-run productivity boost and its long-run learning effect can point in opposite directions, decided by how the tool is used — the exact shape of Karpathy's thesis, and consistent with Strömberg et al. (2026) (AI raises homework, lowers exams, losses concentrated among outsourcers), Bastani et al. (2025) (base GPT-4: −0.19 SD later), and Shen & Tamkin (2026) (engineers who learned a library via AI scored worse unassisted). The meta-analysis (Figure 5, grand mean 0.18 SD across 22 estimates) sorts the whole literature along this axis: losses cluster where AI could do the practice for you, gains where AI plays coach.
Mechanisms#
- Time reallocated, not saved. No effect on total learning time (unlike the large workplace time-savings of Noy & Zhang 2023), but the mix shifts −5.3pp writing / +4.4pp reading-and-searching — effort moves from producing text to absorbing it (task execution → stewardship, à la Copilot users spending half their time verifying).
- Enjoyment up 13% (+0.66 pts, p=0.031): AI made learning more engaging, not mechanical.
- Cheating up but not the driver: rule violations rose +12.6pp combined (p=0.005), yet a back-of-envelope bound attributes at most ~2.2pp (about a third) of the test-score gain to cheating — the effect is mostly real learning.
Beliefs: experience corrects the magnitude, not the direction#
Both groups expect AI to help, but only the treated gauge how much correctly. Control students predict a +25.2pp own gain — about 5× the actual 5.1pp; treated students, who used it, predict +3.8pp, close to the estimate. Perceived gains track actual gains across subgroups. And students already hold the right model: 69% name both a help and a harm channel, 54% say the effect "depends on how AI is used," and their two most-named mechanisms (AI explains/tutors 40%; AI shortcuts the work 42%) map exactly onto augmentation vs automation. Students grasp the mechanism from the outset; firsthand use only calibrates the magnitude — a belief-updating result that rhymes with the professional-misjudgment literature (Becker et al. 2025 / METR: developers predicted +24% speedup, measured −19%).
What it settles — and what it doesn't#
It settles, for this population and task, that off-the-shelf AI can produce durable, objectively measured learning gains, dissolving the "AI must erode learning" prior — but conditions the result on use mode. It does not settle: whether the gains hold outside a proctored lab where time is endogenous; whether they hold for non-elite learners (they skew to the able, suggesting not uniformly); or whether the same augmentation/automation split governs workplace deskilling, where the analogue evidence is still mixed (Brynjolfsson's support agents learn; Budzyń's endoscopists deskill).
Connections#
- The Automation–Optimism Link — the primary complement. The AEI survey found heavy delegators self-report no learning loss; this experiment supplies the objective, randomized measure that survey could not — and it both agrees (augmentation users' learning persists unaided) and exposes the divergence the survey can't see (automation users' gains are hollow, and "automation share" pools both types)
- Outsource Your Thinking, Not Your Understanding — the near-controlled test of the thesis: augmentation (AI helps you understand) builds durable skill; automation (AI does the thinking) leaves nothing once removed — Karpathy's principle, randomized
- AI Brain Fry — both put an objective, measured number on AI's cognitive effect (there, oversight fatigue → +11%/+39% errors; here, learning → +0.27 SD, or hollow gains for automators) against the softer self-report signal; the deskilling half of this paper is that page's mechanism in a learning task
- Returns to Expertise in Agentic Coding — the heterogeneity rhymes: gains skew to higher-ability students, and augmentation ≈ using AI to deepen the understanding that Anthropic's study finds is what amplifies an agent; both say the benefit accrues to whoever brings (or builds) understanding
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the belief half: like reported/anticipated exposure, students' perceptions of AI's effect are directionally right but miscalibrated in magnitude until firsthand use corrects them
- Printing Press Software Democratization — the tension: democratization predicts AI raises the floor, but here gains concentrate in the upper ability quartiles, hinting AI may widen learning gaps rather than close them
- Configurable Human Participation — the system-side mirror, published the same week: HAS-Bench's agency scale runs on the same augmentation-vs-automation axis (A1–A2 automation-oriented vs A3–A5 augmentation-oriented), and both land the same shape of result — the value of AI–human collaboration is decided by its mode and timing, not its amount (there: more agency at A4 breaks tasks A3 solved; here: automation-mode use yields hollow gains that vanish with the tool)
Open Questions#
- Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
- Gains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
- The augmentation/automation choice is endogenous to incentives (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half?
- Does the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer?
Sources#
- Experimental Evidence on the Learning Impact of Generative AI — Contractor & Reyes, Experimental Evidence on the Learning Impact of Generative AI (arXiv 2607.08849, 2026-07-09),
empirical. §4 immediate effects, §5 retention + augmentation/automation heterogeneity, §6 beliefs, §7 conclusion; Figures 4–8, Tables 4–8.
Cited by 9
- AI Brain Fry
Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognit…
- The Automation–Optimism Link
AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-qualit…
- Configurable Human Participation
HAS-Bench (Wu et al., arXiv 2607.04329): a benchmark that makes human participation a *configurable* variable — humans…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated
Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoreti…
- AI Economics & Labor
Map of Content for the ai-economics-and-labor domain — 14 concepts. AI's measured economic footprint: usage telemetry,…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Outsource Your Thinking, Not Your Understanding
"You can outsource your thinking but not your understanding"; understanding as the non-delegable human bottleneck; know…
- Printing Press Software Democratization
Boris Cherny's analogy: 1400s literacy expansion → AI software-writing expansion; domain knowledge displaces coding ski…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
Related articles
- The Automation–Optimism Link
AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-qualit…
- Organizational Complements to AI
The general-purpose-technology argument that AI's productivity gains depend on complementary workflow/skill/org-design…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Unknowns as the Agentic Bottleneck
Thariq Shihipar's map-vs-territory thesis: the gap between what you told the agent and what the work actually requires…
- AI Brain Fry
Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognit…
