H
Howardism
Howardism · Vol. 03Plate II · No. 02

AI Economics & Labor, in order.

Notes14DomainAI Economics & LaborOpen Qs28Newest21 Jul 2026Oldest8 May 2026

Work, wages, org design, and the economics of AI-driven labor.

Map of Content for the ai-economics-and-labor domain — 14 concepts. AI's measured economic footprint: usage telemetry, labor-market effects, returns to expertise, organizational complements, and framing effects on accountability. Curated entry point; see Home for all domains.

  • AI Brain Fry — Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognitive cost surface for both tool and employee framings
  • AI Employee Framing — Kropp et al. (HBR May 2026, n=1,261): framing AI agents as "employees" vs "tools" cuts personal accountability −9pp, increases escalation +44%, reduces error catching −18%, no adoption gain
  • AI Usage Cadences — AEI Cadences report: continuous hourly telemetry reveals AI usage carries the rhythms of daily life — personal use spikes 35%→~50% on weekends, recipes 2.3× at 6pm, sleep advice pre-dawn, tax queries 8× around the Apr-15 deadline; off-hours work skews toward higher-wage occupations
  • The Automation–Optimism Link — AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-quality dimensions (pay, security, job-finding, meaning, autonomy, human interaction), report their skills growing more valuable, and show no learning deficit — inverting the common delegation→deskilling-anxiety narrative
  • Context Advantage, Not Taste — Andrew Ng's reframing of the residual human contribution: not 'taste' but an information asymmetry — 'so long as the human knows something the AI does not, human-in-the-loop is needed.' Recasts the wiki's central open question (is taste a ceiling or the next jagged valley?) as a category error, and makes the human role a closable engineering gap rather than a moat
  • Conversation Artifacts — AEI Cadences report: the 'artifact' (the primary output a user takes away) as a new unit of economic analysis — 93% of conversations produce one, artifact type predicts work/personal/coursework use, compute (tokens) scales with the artifact's economic value, and Claude's output sits ~1 education-year above the prompt
  • Conversation-to-Delegation Shift — OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production'), measured by Codex's share of output tokens across three populations — 99.8% OpenAI / 63.3% organizational / 16.5% individual — with adoption spreading beyond developers; standard usage metrics (active users, chats) become less informative as the unit shifts from a conversation to a delegated workflow
  • Experimental Learning Impact of Generative AI — Contractor & Reyes (arXiv 2607.08849): a randomized, proctored experiment with 211 undergraduates finds off-the-shelf AI access raises immediate test scores +0.27 SD, ~76% of which persists a week later on unaided tests, and lifts essay quality only after AI is removed — but the durable gains belong almost entirely to 'augmentation' users (AI as tutor/explainer) while 'automation' users' (AI-drafts-the-text) short-run gains vanish once AI is gone; the objective, measured-skill counterpart to the AEI self-report that learning both persists and can be hollow depending on use mode
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoretical exposure (tasks an LLM could do), reported exposure (what workers say AI can do today), and anticipated exposure (what they expect in 12 months) — plus their orderings (theoretical > reported > observed) and the GDP, experience, and automation gradients the AEI survey reveals
  • Firm AI-Spend Intensity and Headcount Growth — Ramp × Revelio (Kharazian, Simon & Stevens, June 2026): linking observed firm-level AI-vendor spend (corporate-card / bill-pay) to workforce records for 21,559 US firms, high-intensity AI adopters grow total headcount ~10% and entry-level ~12% over the 24 months after adoption while low-intensity adopters show no significant change — an intensity-gated, learning-curve effect (gains emerge at 6–12 months and compound), broad across roles, but concentrated in the Information sector and drawn from a heavily selected adopter population; a novel spend-side adoption instrument that counters the broad-job-loss narrative
  • Human-AI Accountability Redesign — HBR five-pillar prescription: span-of-control redesign, role redesign, performance management reset, decision-rights/escalation/consequences, agentic-unit-not-human-role design
  • Market-Priced AI Exposure (the AI Premium) — Borri-Liu-Tsyvinski (arXiv 2606.30583): a market-implied AI-exposure paradigm built from 380T tokens of realized AI consumption across 400+ LLMs on OpenRouter, not surveys or task-mapping. An AI Factor (PC1 of token/dollar/user growth) → rolling firm-level AI Betas → a priced AI Premium: a value-weighted long-short earns 64.1 bps/week, concentrated on the intensive/frontier margin (closed-source models, paid/seasoned users, long prompts) and absent on casual/open-weight use; present in developed markets but absent in emerging markets incl. China; the market-implied skill map loads positively on interactive/communication/hands-on work and negatively on analytical/scientific/operations-control (interaction+communication +0.36 SD, Science the single most negative), orthogonal to prior task-based exposure measures (<2% of variance explained); plus early evidence of an agentic economy — tool-call tokens rise from ~0 to 52% of consumption
  • Organizational Complements to AI — The general-purpose-technology argument that AI's productivity gains depend on complementary workflow/skill/org-design changes, not just model capability — David (1990)'s electrification analogy (factories gained only after redesigning around electric motors) and Brynjolfsson's productivity paradox; OpenAI's Codex study supplies the natural experiment: the same model yields 99.8% vs 63.3% vs 16.5% usage across populations, so the gap must be complements (access, permissions, skills, review processes), and digital production may let those complements diffuse faster than electrification did
  • Returns to Expertise in Agentic Coding — Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the actions and 5× the output per prompt, reach verified success ~2× as often, and abandon stuck sessions far less; every occupation lands within 7pp of software engineers; gains are concentrated novice→intermediate, with mastery adding little

Open questions 28 open

  • AI Usage Cadences
    • Time-of-day rests on IP-inferred location; how much noise do VPNs, travel, and datacenter-routed API traffic inject into the "sleep advice pre-dawn" style claims?
    • The weekend personal-use spike is largest in high-income countries — is that a genuine work/life boundary difference, or a composition effect (who uses Claude for what, where)?
  • Context Advantage, Not Taste
    • **Does the asymmetry regenerate faster than it transfers?** The whole human role, under this frame, rests on the answer. Nobody in the corpus has posed it.
    • Ng prefers the frame because it "gives us a clearer path to helping AI systems get better." That is a reason to *adopt* the frame, not evidence that it's true. What would distinguish a context asymmetry from a capability gap empirically? ([[returns-to-expertise]] is the closest thing to an instrument.)
    • If the human's contribution is context injection, is the human *replaceable by better context plumbing* — memory, retrieval, continuous production telemetry — rather than by a better model? That would put the expiry of human-in-the-loop on the infrastructure roadmap, not the scaling curve.
    • Ng writes from 0-to-1 consumer products. Does the frame survive contact with domains where the missing thing is a concept rather than a fact?
  • Conversation Artifacts
    • Tokens are a proxy for both compute cost and output value, but verbose models inflate tokens per unit of intent (the same critique [[conversation-to-delegation-shift]] raises); how much of "compute tracks value" is genuine value vs. models simply emitting more?
    • The reading-level "+1 year" gap may be register (terse prompts, polished replies) rather than substance; can it be separated from genuine elevation of content?
    • Artifact classification is first-party and single-model-graded; do the 30+ categories and the work/personal/coursework split survive independent replication?
  • Conversation-to-Delegation Shift
    • The token-share metric rewards *verbose* agentic output. How much of the 99.8% / 63.3% / 16.5% spread is a genuine work shift vs. agentic tools simply emitting more tokens per unit of human intent?
    • OpenAI-internal is a frontier preview *by assumption*. Does the external organizational curve actually trace the OpenAI path (the paper's implicit claim), or does it plateau where adoption frictions don't vanish?
    • "Asking is half of ChatGPT, doing is most of Codex" — but the two tools self-select different work. How much of the asking→doing contrast is the shift itself vs. routing pre-existing "doing" tasks to the tool built for them?
  • Experimental Learning Impact of Generative AI
    • Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
    • Gains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
    • The augmentation/automation choice is *endogenous to incentives* (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half?
    • Does the same use-mode split govern **workplace** skill accumulation (the open question [[automation-optimism-link]] and [[ai-brain-fry]] leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer?
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated
    • Binned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the report checks robustness with a ≥60%-of-tasks indicator, but the levels remain self-reported)?
    • Reported exposure exceeds observed partly because the survey reaches heavy users; what does the reported/observed gap look like in a representative sample? **Sharpened:** [[market-priced-ai-exposure|realized-consumption measurement (Borri-Liu-Tsyvinski)]] adds a market-implied instrument built on 380T tokens of *actual* paid requests — but it is skewed toward developers/sophisticated users (OpenRouter is ~2% of global tokens), a *different* non-representativeness than the survey's heavy-user skew. The lesson: no current AI-exposure instrument is representative; each collection mechanism biases in its own direction, so the reported/observed/market-implied gaps are partly artifacts of *who each method reaches*. A representative census remains the open target.
  • Firm AI-Spend Intensity and Headcount Growth
    • **What operational mechanism converts intensive AI spend into hiring?** The paper establishes the correlation (adopters, especially intensive ones, grow) but explicitly cannot say *why* — product acceleration, sales productivity, engineering leverage, support automation, faster analysis, or new business lines are all candidates, and the firms that cracked it have no incentive to share.
  • Market-Priced AI Exposure (the AI Premium)
    • **How much does the developer skew move the answer?** OpenRouter's slice is unrepresentative; would a representative realized-consumption panel (if one existed) price the same firms and skills, or is the frontier/intensive-margin concentration partly a sampling artifact of who uses OpenRouter?
    • **Why is the market-implied skill map orthogonal to every task-based measure (<2% variance)?** Is market-implied exposure capturing genuinely different information (forward-looking rents, complement/substitute *value* rather than technical automability), or is it noisier — and which should labor-impact forecasts trust?
  • Organizational Complements to AI
    • The "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign workflows quickly, or does the low cost of *tool* adoption mask slow, expensive *process* redesign (the real complement)?
    • Which complement is the true binding constraint — access/permissions, skills, or review capacity? The paper lists all; it doesn't decompose their relative weight.
    • If complements, not capability, gate value, does model progress have *diminishing* near-term returns until orgs catch up — and how long is that lag for agentic AI specifically?
  • Returns to Expertise in Agentic Coding
    • Outcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the whole success gradient — is *real* outcome vs. who-narrates-success-in-the-transcript?
    • The study excludes headless / SDK / IDE usage (a "substantial share"). Does the returns-to-expertise pattern hold in non-interactive and pipeline use, where there is no human steering mid-session at all?
  • The Automation–Optimism Link
    • Selection vs. treatment: tenure controls attenuate but don't eliminate the enthusiast-selects-into-delegation story. Does a within-person design (sentiment before/after adopting automated workflows) hold the effect?
    • The sample is heavily computer/math + management and 88% men; how much of the automation–optimism link survives in a representative population?