H
Howardism
Plate IIAI Economics & Labor中文HOWARDISM

Organizational Complements to AI

The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design changes (David's electrification analogy, Brynjolfsson's paradox) — OpenAI's Codex natural experiment (99.8% vs 16.5% usage of the same model) shows the gap is complements; also home to the HAT substitution model and Kalff & Simbeck's institutional complement.

Article metadata
Publication details
Published:June 26, 2026
Filed:Concept
Domain:AI Economics & Labor
Reading:70 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Organizational Complements to AI

Sources#

Summary#

The economics frame OpenAI's Codex usage study uses to explain why agentic-AI adoption is so uneven across populations under an identical model: a long tradition of general-purpose-technology research holds that productivity gains from a new technology depend on complementary investments — in business processes, worker skills, organizational design, and intangible capital — not on the technology's capability alone. The canonical case is David (1990)'s dynamo: early factories swapped centralized steam engines for centralized electric motors while keeping the old plant layout, and got little. The large gains came only decades later, when firms redesigned production around electricity's distinctive affordance (small, decentralized motors → reorganized factory floors, new task sequencing, flexible layouts). The history's lesson, applied to agentic AI: near-term effects may understate long-run potential, because firms have not yet discovered or scaled the new production processes the technology makes feasible — Brynjolfsson's "modern productivity paradox" restated for AI.

Evidence note. empirical for the Codex cross-population data; the GPT/complements framing is the paper's synthesis of prior economics literature (David 1990; Brynjolfsson, Rock & Syverson 2019; Demirer et al. 2026). It is an interpretation of the adoption gaps, not a measured causal estimate of complements.

The three-population gap as a natural experiment#

The paper's sharpest empirical move: if adoption depended only on model capability, usage would look similar wherever the same model is available. It does not. Codex's output-token share is 99.8% (OpenAI) / 63.3% (organizational) / 16.5% (individual) — and OpenAI workers use it across far more job functions and at far higher concurrency. Since the model is constant, the gap must be complements:

  • Access to relevant files, repositories, and systems
  • Permissions and security requirements
  • Workforce skills and familiarity with frontier models
  • Management expectations and organizational buy-in
  • Complementary review processes for verifying delegated work

OpenAI is the high-complement extreme (cheap marginal usage, training campaigns, feedback loops, model-adjacent workflows), which is exactly why its usage is a frontier preview rather than a population estimate. The conclusion: agentic AI is not simply a cheaper input into existing work — its value depends on whether organizations can redesign workflows, responsibilities, and review processes around delegation and verification.

Why this transition may be faster than electrification#

The paper flags one disanalogy that cuts the other way. Electrification required firms to redesign physical plants and replace durable capital — slow and expensive. Agentic AI lets workers and organizations experiment with new workflows at low cost: no factory to rebuild, just process and tooling to rearrange. This lower cost of experimentation may let new production methods diffuse faster than in prior GPT transitions — even though the full organizational complements are still emerging. The within-OpenAI evidence supports speed: between Dec 2025 and April 2026, later-adopting functions (legal, recruiting) went from ~0 to ~75% Codex token share, with the steepest stretch ~20%→75% in a single month, riding an internal adoption campaign.

The urgency claim that inverts this page's analogy ("We Must Act Now", July 2026)#

"We Must Act Now: A Statement on AI's Transformation of the Economy" (July 13, 2026) is the first time the vault carries this page's thesis as a policy prescription rather than as an explanation. Brynjolfsson — whose productivity paradox is the framing at the top of this page — organized it with Ajay Agrawal, Anton Korinek, and Tom Cunningham, and states the normative form directly: "guide AI to complement humans rather than simply imitate them." Over 200 economists and AI researchers signed, sixteen of them Nobel laureates.

The interesting part is that its central urgency claim runs against the analogy this page is built on. Korinek: "Steam, electricity, and computers each gave societies decades to adapt; AI may give us only a few years. We cannot improvise our strategy and institutions in the middle of the transformation." David (1990)'s point about electrification is that the decades were required — the gains arrived only after firms rebuilt production around the new affordance, and could not have arrived sooner. The letter's premise is that the same adaptation must now happen inside a window an order of magnitude shorter. Those are not compatible as stated; the letter asserts the compressed timeline and never argues for it.

Evidence note. practitioner-opinion — a news release announcing an open letter, 558 words, no measurement of any kind. Signatory count is not evidence: 200 economists asserting a timeline is one claim, not 200. The section above ("Why this transition may be faster than electrification") is the only argued version of the compressed-timeline case in the vault, and it rests on one favorable internal case at OpenAI.

Everything measured on this page points the other way — at lag, not speed. Emergence Capital's RPE gap, Ramp × Revelio's intensity gate and 6–12-month learning curve, ICONIQ's enablement and governance overruns, and Kalff & Simbeck's German firms whose advanced analytics stall on uncentralised data are four independent instruments finding that complements are slow and mostly not yet built. The single vault datum for speed is the within-OpenAI adoption campaign (~0 → 75% Codex token share in months) — the most complement-rich organization on earth, i.e. the least representative case available.

Which leaves the letter's actual argument in better shape than its headline. If complements are the binding constraint and they take years to build, "begin now" follows without needing the transformation to be fast — the case for acting early is strongest precisely when adaptation is slow. The letter's own quotes concede the uncertainty its headline elides (Spence: "a high level of uncertainty about the magnitude and timing"; Cunningham: "we are driving in the fog").

The complement that binds: verification and coordination#

Across the literature the paper cites, the recurring complement is supervision/verification/coordination capacity. Hitzig et al. (2026) (which this paper cites) argues agentic systems move interaction from assistance toward delegation, "making supervision, verification, and coordination central determinants of value creation while increasing returns to domain expertise." Demirer et al. (2026b) show large task-level gains translate only imperfectly into output because downstream human activities remain bottlenecks; Demirer et al. (2026a) find AI helps most when it can execute contiguous chains of tasks (workflow adjacency matters). The through-line: the missing complement is usually not a better model but a redesigned review-and-coordination process around the delegated work — the org-level form of Verification as the New Bottleneck.

External corroboration: the AI revenue-per-employee lag (Emergence Capital, June 2026)#

The Codex study makes the complements argument on usage telemetry; Emergence Capital's Beyond Benchmarks 2026 makes the same argument on company financials, from a different direction. Across 50K+ operating companies (Standard Metrics financial-benchmark cohort), AI companies generate ~39% less revenue per employee than non-AI companies in every segment (the AI-vs-non-AI split is published at the top decile only — the ~39% is the average of the four top-decile band gaps) — precisely the "gains lag adoption" pattern this page predicts. The report's own framing is a near-verbatim restatement of the complements-lag thesis: "AI is not yet a shortcut to best-in-class efficiency, it's an investment phase… expect a lag between AI adoption and measurable gains in revenue per employee as companies scale usage and translate capability into output." The diffusion is visible in the same data: AI-native RPE is growing faster than non-AI (up to +58% YoY at the $100M+ top decile while non-AI declined −6%), consistent with complements being built and the gap closing over time. This corroborates the David-electrification claim with cap-table-grounded financial metrics rather than usage shares — with the caveat that it is a VC-published (though data-partner-sourced) dataset. The startup-side treatment lives at AI Investment Story, Not Efficiency Story.

Firm-level corroboration: the AI-spend intensity threshold (Ramp × Revelio, June 2026)#

The Codex study argues complements from usage shares and Emergence from financials; Kharazian, Simon & Stevens supply the same argument from firm-level adoption spend, and it is almost a direct measurement of the thesis. Linking Ramp AI-vendor payments to Revelio workforce records for 21,559 US firms, they find the employment gains from AI are gated by intensity and delayed: only high-intensity adopters (~$34/employee/month) grow headcount (10% over 24 months), while low-intensity adopters ($2.78/employee/month — the enterprise-chat-subscription tier) show no detectable change. The conclusion is a near-verbatim restatement of the complements-lag: "Enterprise chat subscriptions do not appear to be enough. Nor are a few months of experimental spending… benefits require complementary investments, organizational change, and learning inside the firm. Many firms may buy subscriptions, run pilots, and then fail to make the sustained investments required to benefit." The 6–12-month lag before gains appear and their compounding over 24 months is the David-electrification learning curve observed on payment traces. The full treatment lives at Firm AI-Spend Intensity and Headcount Growth; the sector concentration (gains significant only in Information) is itself a complements story — the complements are furthest developed where coding-agent workflows already exist.

The complements, itemized on a builder's own P&L (ICONIQ, Q2 2026)#

Where the Codex study infers complements from an adoption gap and Ramp measures a spending threshold, ICONIQ's State of AI 2026 survey (~305 AI-building software companies, empirical) prices the complements out as line items — and its respondents' own explanation of why internal AI is expensive is a near-verbatim complements list:

  • Internal AI-systems spend is projected to jump from a prior 1–3% of revenue to 11%, then a projected 16% in 2026 — and the figure is defined broadly on purpose, to capture the "true cost of AI" beyond tokens.
  • Respondents say true cost is hard to predict, and the overruns come from exactly the complements this page names: (1) token spend scaling non-linearly once single calls become multi-step agentic pipelines (a $0.10/run workflow reaching $1.50+ on retries); (2) data infrastructure — production RAG, permissioning, structuring; (3) organizational enablement — governance frameworks, usage standards, sustained training, "costs that rarely appear in initial business cases." Items (2) and (3) are the intangible/process complements (Brynjolfsson) restated as budget-variance sources.
  • The productivity payoff is present but gated by these complements: agentic tools specifically return <30% gains across every revenue band (vs coding assistance at ~48% for high-growth) and "often require human intervention" — the verification/coordination complement binding again. The internal-productivity-as-a-moat exemplar (Ramp: 350+ Git-versioned reusable workflows) is a complement being built deliberately. Full unit-economics treatment at AI Product Economics Maturation.

The caveat: this is a self-reported survey with prediction-grade forward figures, so the 16% projection is intent, not measured spend — but the decomposition of where cost surprises come from is respondents describing complements they underinvested in.

Randomized corroboration: the complement that didn't get built (Jabarian & Henkel, July 2026)#

Every argument above is inferred — from an adoption gap, from financial ratios, from a spending threshold, from budget-variance anecdotes. Jabarian & Henkel's hiring field experiment supplies the same thesis under randomization, and it is unusually clean because the missing complement is visible as a cost in the same dataset as the capability gain.

The firm automated one stage — the job interview — and left evaluation with human recruiters. The automated stage worked: 12% more offers, ~18% more job starts and one-month retention, no productivity decline. The process around it did not adapt, and two measurements show it:

  • The evaluation queue lengthened 2.8×. Median interview→offer-decision time went from 2.62 days (human-led) to 7.24 days (AI-led), because recruiters now review conversations they did not have. Scheduling time fell (0.51 → 0.32 days), but not nearly enough to compensate: end-to-end time-to-hire rose from 20 to 24 days (p=0.033). The firm bought a better funnel and paid for it in latency at the one stage it did not redesign — the Verification as the New Bottleneck complement, measured causally rather than inferred.
  • Humans discount the AI's signal. Recruiters rate AI-conducted interviews higher (mean score 1.90 → 2.01) yet weight them less in the offer decision: the interview score is significantly less predictive of offers in the AI arm (interaction −0.047, −0.029 with controls) while the independent language-test score becomes more predictive. The discount is significant only among recruiters who told the survey they consider interview performance more important than test scores — precisely the people whose decision rule the automation displaced.

The paper's conclusion is a near-verbatim statement of this page's thesis, arrived at from a randomized design: "Firms can increase the efficiency of screening by interviews through automated standardization, but realizing the full returns from automation requires complementary adaptation in how humans rely on AI signals for decision-making."

Two things this adds beyond corroboration. First, it names a complement the page's other sources do not: trust calibration in the human evaluator — not access, permissions, skills, or review capacity, but whether the downstream human weights the machine's output correctly. Second, it is a counterexample to the usual framing that complements gate whether gains appear at all. Here the gains appeared immediately and at scale; the missing complement showed up as a partly offsetting cost on a different margin (latency, discounted signal). Complements do not only determine whether value is captured — they determine where the bill arrives.

The complement that isn't on the list: institutions (Kalff & Simbeck, July 2026)#

Kalff & Simbeck's mixed-methods study of German HRM (arXiv 2607.13839, empirical) is this page's argument observed on a national sample of one business function — and it adds a kind of complement the itemized list above does not contain.

What it is. Three instruments on one question: 14 semi-structured expert interviews (AI experts, tool vendors, HR managers, civil-society NGOs), three group discussions with advisers to German works and staff councils (Betriebsräte), and a survey of 427 HR managers fielded February 2025 by an ISO-20252-certified provider, 410 valid after removing implausible entries. Analysis is descriptive; no causal claim is made.

The demand side of the complements argument, ranked. Respondents ranked their employers' five most important reasons for adopting AI (5 points for first place down to 1 for fifth, cumulated). Summing the paper's own four clusters:

Cluster (Figure 2)Weighted scoreLeading items
Rationalisation2,887increase process efficiency 1,138 · save costs 917 · automate routines 562
Human Capital1,120addressing skill shortage 311 · increase creativity 298
Analysis and Control1,081improve decisions 367 · increase process objectivity 267
Prediction716increase planability 382 · plan work organisation 272

Rationalisation is more than double any other cluster. The two lowest individual items out of seventeen are "improve value creation" (102) — sitting inside the rationalisation cluster — and "predict events" (62), dead last. The paper's conclusion follows directly: firms bought AI to do the existing work cheaper, and "the actual promises of augmentation through HR analytics, such as in-depth analytical insights, predictions or prescriptive action, play no role in the actual practice of enhanced HR work."

Adoption is uneven in exactly the shape complements predict. Figure 1, share of the 410 reporting each tool: Personalbedarfsermittlung (workforce-needs planning) 38%, Bewerber:innen-Vorauswahl (applicant pre-selection) 35.6%, Karriereseiten und Stellenanzeigen (career pages and job ads) 26.3%, Lern- und Entwicklungsbedarf (learning-and-development needs) 23.7%, Leistungsbewertung (performance evaluation) 21.5% — and 20.2% answer Nichts davon, none of these: one German HR department in five using no AI tool at all, three years after ChatGPT. The blockers the interviews name are this page's list nearly verbatim — fragmented, non-centralised data ("independent HR systems at each regional office or subsidiary"), firm size (SMEs "fail to see immediate returns from complex AI-driven applications, which typically depend on large volumes of centralised data and stable, standardised processes"), digital maturity, and managerial support.

The complement the list is missing. Access, permissions, skills, review capacity, trust calibration — every complement itemized above is internal to the firm. Germany supplies a fifth kind, and it is the paper's own claimed contribution: institutional. The Works Constitution Act (Betriebsverfassungsgesetz) §87(1) no. 6 subjects any technology capable of monitoring employee performance or behaviour to works-council co-determination, and the EU AI Act puts candidate screening and predictive performance management in the high-risk tier. What the interviews show is that this does not merely slow adoption — it steers which capability gets bought, through three observed channels:

  • Capability substitution. "Some organisations adopt simpler chatbots or generative-text assistants to sidestep these compliance obligations." A vendor (DEV3) deliberately engineers personal data out of its ML demand-forecasting product: "we don't even have personal data, so we don't have a GDPR or DSGVO thing and we don't have an AI Act thing either… It's just aggregated data, there's not even a single incident where a person plays a role."
  • Jurisdictional arbitrage. Globally active firms outsource HR functions to affiliates and global service centres, removing them from local works-council jurisdiction — "increasingly common when organisations wish to avoid negotiations regarding sensitive AI-based analytics" (COD1). The firm moves the function rather than forgoing the technology.
  • Label management. Vendors play the "AI" label up to support a business case; firms play it down to avoid co-determination scrutiny. Treated at AI Employee Framing.

The synthesis the paper doesn't state: the predictive tools that survived are the ones with no person in them. Reconciling both figures against §4.2's regulatory prose, the adopted predictive use cases are aggregate and impersonal — workforce-needs planning is the most-adopted tool at 38%, and "increase planability" (382) and "plan work organisation" (272) are the two high items in the Prediction cluster. The predictive tools pointed at individuals sit at the bottom of Figure 1: turnover prediction (Fluktuationsvorhersage) 11.7%, culture analysis (Kulturanalyse) 8%, sentiment analysis (Sentimentanalyse) 4.6% — with "predict events" last in Figure 2 at 62. That line falls exactly where §87(1) no. 6 and the AI Act's high-risk tier fall. This reconciliation across Figures 1 and 2 and §4.2 is the wiki's, not a claim the authors make.

The precedent, in the same country. The authors note the pattern "reflects earlier IoT and Industry 4.0 initiatives in Germany, which also focused mainly on efficiency gains" (Kalff 2019; Butollo, Jürgens & Krzywdzinski 2019). That is a completed transition in the same institutional setting that resolved to rationalisation rather than redesign — a less flattering companion to the electrification analogy at the top of this page, where the redesign eventually arrived.

Evidence note. empirical, and the weakest instrument on this page: cross-sectional, single-country, and self-reported for every quantitative measure — a limitation the paper states itself. Two further cautions. The project (TranKI) is funded by the Hans Böckler Foundation, the research foundation of the German trade-union confederation, and "AI as rationalisation, not augmentation" is the labour-side reading of this evidence; the authors declare no conflict and present the qualitative material in both directions, but the framing is not funder-neutral. And the survey's central construct is unstable in respondents' own heads — see Telemetry vs. Survey Measurement for why that undercuts the "predictive analytics plays no role" finding specifically.

The formal model of this thesis, and where it puts its assumptions (Banerjee & Singh, July 2026)#

Banerjee & Singh's Human-AI Task Allocation (HAT) model (arXiv 2607.20781) opens with this page's argument stated as its motivating premise: "Organizations do not replace employees simply because AI becomes technically capable of performing a task. Replacement is an organizational decision that depends simultaneously on expected costs, risks, organizational structure, managerial coordination, and the relative economic characteristics of human and AI labor." It then builds a hierarchical optimization around that premise. What it contributes is vocabulary and a parameter list — not evidence.

Evidence note. practitioner-opinion — a formal model with no empirical data, the same tier and the same shape as The Tragedy of the Cognitive Commons. Every result is a theorem conditional on stated assumptions; the paper is disciplined about saying so ("the results should be interpreted as conditional statements… The purpose of the HAT framework is not to claim universal inevitability of AI substitution"). Where the vault holds a measurement bearing on one of its predictions, the measurement wins and is named below.

The machinery, compressed. An organization has depth D managerial layers over line workers at level D+1, with level sizes derived from spans of control (e_i = ∏ s_ℓ, after Garicano 2000). A human line worker costs Min$ + ½·Δ$·D·μ (compensation rising in skill μ); a manager costs C_{D+1}·(1+r_0 D)^{D-i+1}, so human cost escalates super-exponentially toward the top. An AI agent costs T_k/n_k + M'_k + ½·Δ$'·D·μ' — a fixed training cost amortized over n_k deployments, plus a marginal operating cost, plus a capability term. Both sides are risk-adjusted at the level of the individual agent: C̃ = C + λR, where λ is the firm's risk sensitivity and AI risk decomposes as R' = ω₁R^(rel) + ω₂R^(comp) + ω₃R^(rep) — reliability, compliance, reputation.

Assumption 1 ("Human-AI Cost Asymmetry") carries the whole paper: (i) Δ$' ≤ Δ$, AI capability costs grow no faster than human skill costs; (ii) T_k/n_k → 0 as deployments scale, so the AI baseline approaches a capability-independent floor M'_k; (iii) r'_0 ≤ r_0, AI coordination cost escalates no faster with depth than human coordination cost. The named result, the Human-AI Substitution Principle (Theorem 5), is then one line: replace the human iff C'_ik + λR'_ik < C_ij + λR_ij. Everything else — abrupt transitions (Thms 7, 14), flattening (Thm 4, Cor. 1), middle-management vulnerability (Cor. 2), hybrid organizations (Thms 10–11), irreversibility (Cor. 5), an upskilling-vs-AI-investment Nash equilibrium (Thm 15) — is derived from that comparison plus Assumption 1.

The complements are inside r'_0, and the model assumes them away. Table 1's own observable gloss for r'_0 is: "Integration effort, workflow orchestration, monitoring, human oversight requirements." That is this page's complements list, verbatim, compressed into one scalar — and Assumption 1(iii) asserts it is no larger than the human coordination multiplier. The paper concedes the assumption is unfalsifiable as stated: r'_0 "is not directly observable and should be calibrated as a scenario parameter," bracketed between 0 (fully scalable AI coordination) and r_0 (human-equivalent). The vault has the measurement it declines to bracket above r_0. In Jabarian & Henkel's experiment, automating the information-collection stage made the remaining human stage 2.8× slower (interview→offer 2.62 → 7.24 median days) and end-to-end time-to-hire longer (20 → 24 days, p=0.033); oversight capacity does not expand when output does. A single measured case is not a refutation, but it is one instance of r'_0 > r_0 on the margin the model needs it small.

The seven predictions, against what the vault measures#

Table 3 is the paper's most useful part — falsifiable predictions with named empirical signatures. Four of the seven can be checked against evidence already here.

Prediction (Table 3)HAT's stated signatureWhat the vault has
P1 Discontinuous automation near thresholds (Thms 7, 14)sharp phase transitions, not smooth S-curves; workforce change clusters at capability/risk-reduction eventsSplit, and it splits on which margin. Adoption does look abrupt: inside OpenAI, later functions (legal, recruiting) went ~0 → ~75% Codex token share with the steepest stretch 20%→75% in a single month. Workforce adjustment does not: the Ramp × Revelio panel at Firm AI-Spend Intensity and Headcount Growth has its high-intensity event study rising 0.003 → 0.020 → 0.071 → 0.188 → 0.277 → 0.452 across months 0/3/6/12/18/24 — a compounding ramp on a 6–12-month learning curve. P1 predicts the second, and the second is the one that looks gradual. A third margin now cuts the other way: Matsui's 300M-work OpenAlex panel finds the composition of a knowledge-production team breaking at the capability event itself — a decades-long decline in solo authorship halting or reversing in 23 of 26 fields, dated to ChatGPT's release, with the decline steepening to −2.9 pp in 2022 and then stopping dead (+0.1 in 2023). That is the closest thing the vault has to "workforce change clustering at a capability event," and it is the outcome P1 actually names. Two discounts: it is a kink (slope change) rather than the level jump a phase transition implies, and roughly half the pooled break is venue composition (+1.72 → +0.75 pp/yr inside continuously observed venues)
P2 Middle-management vulnerability (Cor. 2)AI reduces middle layers before top or bottom; coordination roles automated firstNothing supports it; the closest measurements point at the bottom. At high-intensity adopters, manager-plus headcount grew +6.5% while entry-level grew fastest (+12.0%); only the manager-plus share fell (−1.52pp). Brynjolfsson/Chandar/Chen (via The Tragedy of the Cognitive Commons) find 22–25-year-olds in the most AI-exposed occupations −16% while 35–49-year-olds in the same occupations +8% — mid-career growth, junior decline, the reverse ordering. German HR is a third case pointing the same way: the only displacement Kalff & Simbeck report is clerical — a chatbot absorbed list-merging work and "employees who previously performed these duties were laid off" (COM1) — and the named casualty is "career paths for entry-level HR professionals," not a management layer. See the caveat below on where Cor. 2 comes from
P3 Organizational flattening (Thm 4, Cor. 1)fewer hierarchical levels; wider spans of controlThe one prediction with confirming evidence. ICONIQ (via AI-Native Organization): at $100M+ scale, 72% of companies with 50%+ AI revenue run on 1–4 management layers vs 56% of peers, and high-growth firms are widening spans — first-line R&D managers with 7+ reports 21%→30%, GTM 26%→33% (2025→2026). Exactly the signature. But it is cross-sectional self-report from AI builders, so selection is uncontrolled, and the widening runs against the oversight-capacity ceiling the pages above measure. German HR supplies a mechanism running the other way: works-council advisers and HR leads (COD1, COM4) report tightly integrated HR analytics being read as "overly rigid forms of micro-management," clashing with the self-organisation and empowerment agenda, with predictive workforce analytics threatening to "reintroduce hierarchy quantification into daily practices." On that account flatness is a management philosophy the analytics collide with, not a consequence they produce — the opposite causal direction from Thm 4
P4 Regulatory persistence of hybrid structures (Thms 10, 11)regulated industries retain humans longer; substitution starts in low-compliance roles; certification events trigger adoption spikesFirst field evidence, direction confirmed, mechanism richer than the prediction. Kalff & Simbeck (section above) observe substitution starting in the low-compliance corner: firms "adopt simpler chatbots or generative-text assistants to sidestep these compliance obligations," a vendor engineers personal data out of its forecasting product specifically to escape GDPR and the AI Act, and the person-level predictive tools that trigger §87(1) no. 6 co-determination sit at the bottom of reported adoption (turnover prediction 11.7%, culture analysis 8%, sentiment analysis 4.6%) while impersonal planning tools sit at the top (38%). Two channels P4 does not contain: regulation steers which capability is bought, not only how fast; and firms move the function out of jurisdiction (offshoring HR to global service centres) rather than forgo the technology. Caveats: self-reported, cross-sectional, one country, one function — the level-resolved panel in the open question below is what would test it properly. The prior nearest thing was another untested prediction: Lovett's five vulnerability factors put regulatory intensity and safety criticality as protective, predicting software/financial analysis/legal research degrade before medicine and engineering
P5 Asymmetric skill-cost evolution (Asm. 1, Cor. 3)flat-Δ$' occupations substitute fast; counterintuitively, high-skill workers more vulnerable thereThe sign is an artifact of the instrument, and P5 states it as determinate. Steele & Cruz's seven-instrument head-to-head at Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated finds the exposure–salary gradient positive for the four newest instruments and strongly negative for Frey & Osborne 2017 — the whole white-collar-displacement question answered by measurement philosophy. Separately, Returns to Expertise in Agentic Coding finds expertise amplifies the agent (2× actions, 5× output per prompt, ~2× verified success), a complementarity the model's linear objective cannot represent at all (its own limitation 3)
P6 Deployment-scale acceleration (Min$' = T_k/n_k + M'_k)large firms hit thresholds first; industry adoption cascades; first-mover structural advantagePartly consistent, but the vault's better-identified variable is different. Ramp × Revelio's adopters are larger, more technical, higher-paying pre-adoption — consistent with P6 and confounded with selection. The gate that actually separated outcomes was per-employee intensity ($2.78 vs $33.67/month) and time, not headcount to amortize over. HAT has n_k but no term for the organizational investment that makes the n_k-th deployment work; the paper's limitation 2 concedes n_k is exogenous. German HR corroborates the size gradient qualitatively (SMEs "fail to see immediate returns from complex AI-driven applications, which typically depend on large volumes of centralised data") while exposing a channel with no n_k at all: 183 of 410 respondents report using AI tools informally on personal devices regardless of employer policy or company-provided systems — adoption happening one level below the entity that does the deploying
P7 Strategic human upskilling (Thm 15)reskilling rises where the risk-adjusted gap narrows; migration toward high-ω₁/ω₂ tasksComplicated by the vault's other framework paper, on the supply side. Theorem 15 makes upskilling a private effort choice u_j ∈ [0, ū_j] against a convex private cost. The Tragedy of the Cognitive Commons argues the developmental experience that produces u_j is a profession-level commons whose regeneration mechanism — entry-level work — is what AI removes, and that no single firm has an incentive to maintain it. The contest formulation has no room for a shared upper bound on ū_j. Experimental Learning Impact of Generative AI supplies the mechanism-level version: upskilling gains persist for augmentation users and vanish for automation users once AI is removed. German HR is directionally consistent but evidentially thin — firms report investing in reskilling and HR managers report needing "dual competencies" (data-analytic plus humanist), all self-reported intent rather than measured effort or outcome

One caveat that travels with P2, the paper's most-quoted result. Theorem 8 proves only that cost-based substitution pressure rises monotonically toward the top — on cost alone the CEO is the most attractive replacement, and the calibrated example shows exactly that ($60.6k of gap at the line-worker level rising to $173.7k at the CEO). The intermediate peak requires a second curve, substitution feasibility F(i), that falls near the top. F(i) is neither derived nor measured: in the calibration it is five numbers set by hand (F = 0.50 / 0.80 / 1.00 / 0.55 / 0.25 from line worker to CEO), chosen so that the product peaks at level 3. Corollary 2 is honest about this — it says middle managers are most exposed if a single-crossing condition holds — but the headline travels without the conditional.

Two structural results the vault's evidence contradicts#

Extreme-point optimality. Theorem 6 puts a linear objective on a simplex, so the optimum sits at a vertex: the whole task goes to the single lowest-risk-adjusted-cost agent, and mixtures arise only from exact cost ties or externally imposed constraints (Thm 11). The vault's only randomized substitution result is the opposite shape — the hiring experiment split one task by stage, automating information collection and leaving evaluation entirely human, and that split is what produced +12% offers and +18% retention. HAT can represent it only as a constraint imposed from outside the optimization, never as the optimum. The paper's limitation 4 concedes the point ("best interpreted as benchmark structural tendencies rather than literal descriptions of operational firms") and limitation 5 concedes the single-task abstraction that rules out stage decomposition.

And the model cannot express a variance advantage at all. This is worth stating plainly, because the vault's best causal evidence says the winning margin was variance. Three separate reasons the cost conditions cannot encode it:

  1. Output quality is assumed identical. §5.1 derives the objective from V = B − C − λR and drops B because "B is fixed for a given task." Both agents produce the same benefit by construction. Controlled variance is a claim about the dispersion of output quality across repetitions — the model has neither repetitions nor variable quality.
  2. R is a level, not a moment. λ is called risk sensitivity and C + λR has the shape of a mean-variance functional, but Table 1 defines R as error/turnover/absenteeism for humans and reliability/compliance/reputational failure for AI — an expected-loss level. No second moment appears anywhere in the paper.
  3. Even a generous encoding collapses the distinction. One could set R' < R to say the AI is more reliable — but that is a mean shift in effective cost, indistinguishable from the AI simply being cheaper. The model cannot separate "AI is better on average" from "AI is more consistent," which is precisely the distinction the experiment identifies: the AI beat only 61%/64% of human recruiters on topic coverage and vocabulary richness while beating 100% on question-guideline similarity and 83% on order adherence. Not more capable — less dispersed.

The direction is not incidental. Every worked example in the paper runs the other way (R'_5k = $20k against R_5j = $5k), and all seven predictions are built on λ(R' − R) > 0 — AI as the riskier party whose risk penalty is what delays substitution. A substitution model in which AI's reliability is the barrier has no way to describe the case where AI's reliability is the product.

The calibration's own headline, and what the vault says happened. §5.7 calibrates a five-level, 341-person professional-services firm (Min$ = $70k, Δ$ = $12k, r_0 = 0.10; Min$' ≈ $9.95k, Δ$' = $0.5k, λ = 1.5, R' = $20k vs R = $5k) and computes the skill threshold at μ* ≈ −1.52. Since skill is non-negative, every line worker sits in the vulnerability zone — the model's verdict is that a mid-sized advisory firm should already have replaced all 256 analysts, driven almost entirely by the $70k-vs-$10k baseline gap. Raising compliance risk to a healthcare-like setting moves it only to μ* ≈ −0.46, still negative. Against that: at intensively AI-adopting US firms, entry-level headcount grew 12.0% over 24 months and entry-level share rose 1.15pp. The model is normative and the measurement is descriptive, so this is not a formal contradiction — but it is the most checkable claim the paper makes, and the vault's best firm-level evidence points the other way.

What the model is genuinely good for, stripped of its predictions: it names the levers cleanly. Substitution decisions decompose into nominal cost, risk sensitivity λ, and risk-profile difference R' − R; the three AI risk components map one-to-one onto technical, social, and legal-regulatory accountability (§5.2, and see Human-AI Accountability Redesign); and governance investments — auditability, regulatory engagement, transparency — become parameters of the substitution boundary rather than obstacles to it. That decomposition is portable whether or not any theorem built on it survives contact with data.

The arithmetic, stated by a practitioner (Ng, August 2026)#

Andrew Ng gives the complements thesis its shortest statement in the corpus, to a general audience (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, 2026-08-28, practitioner-opinion), and attributes the method to Erik Brynjolfsson and Andrew McAfee: economists "have analyzed many people's jobs… break it down into individual tasks and maybe AI could do… 30 40% of many jobs and what that means is well that 60% that a human does has become even more valuable because it's called an economic complement to the 30 40% that's now cheaper." The displacement he does expect is within the occupation — "people that use AI will replace people that don't use AI" — not of it. Two further remarks are this page's argument from the deployer's chair. Asked how he measures AI's productivity effect: "the business outcome of AI is more a function of the business than a function of the AI… the KPIs tend to be related to the business rather than the AI" — the complements gate stated as a measurement problem (there is no AI-side KPI because the value is realized, or not, in the redesigned process). And his account of embedding "recruiting engineers… professional engineers that sit in a recruiting team" is a named complement — the engineer-in-the-function pattern the Codex study's access/skills/review list abstracts over.

Weigh it accordingly. The 30–40% is a theoretical-exposure number (the share of tasks a model could do), and the vault's observed-exposure instruments run well below it — ATLAS's median occupation has AI present in 21% of its tasks (Task Saturation: Broad but Shallow AI Diffusion; the full placement is on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated). The "60% becomes more valuable" step is the complements thesis, and it is exactly the step the measured evidence on this page finds gated rather than automatic: complementary value shows up at high-intensity adopters on a 6–12-month curve and not at all at the subscription tier. Ng's version has no gate.

The same thesis inside the research production function (ATLAS Science, September 2026)#

Every case on this page so far is a firm. Google/DeepMind's ATLAS science report (Google AI & Economy ATLAS: AI in Science (September 2026), empirical with vendor COI) runs the argument in a production function that has no P&L, no CIO and no rollout plan — science — and gets the same shape, which is the most useful thing about it: the complement constraint is not an artifact of corporate procurement.

The measured pattern, from a survey of 637 US/UK scientists (self-reported throughout): just under three quarters save time with AI, averaging 6.9 hours a week, and ~84% report more lab output over three years — a real local gain. But 43.5% say their primary rate-limiting bottleneck moved downstream over two years (against 13.8% upstream), physical experimentation and data collection is now the largest single bottleneck at 24%, 40.5% report a growing backlog of untested hypotheses, and among those saving time 45.7% spend more than a quarter of it verifying AI output. Acceleration of the analysis stage piles work into the stages that cannot be accelerated the same way.

The authors reach for exactly this page's literature to explain it — Kremer's O-ring, Demirer et al. (2026) on task chaining, Gans & Goldfarb (2026), Garicano et al. (2026) on redrawn job boundaries — and conclude that "the residual tasks in the bundle of scientific work are harder to scale and absorb the time savings," so realizing the gains "may require redesigning scientific processes and workflows." That is the electrification analogy with a wet lab in it, and it adds two things the firm cases do not.

  • The binding complement here is physical capital, not process. In the German HR and Codex cases the missing complement is organizational — data centralisation, permissions, redesigned workflow. In science, 21% of the reinvested time dividend goes straight back into physical lab execution and data collection, and the named constraints are wet-lab capacity, clinical validation and field data collection. Those are not reorganizable at software speed at any price a lab controls, which is a harder version of the same claim.
  • It is a sector where the adopters have unusual autonomy. Scientists largely choose their own tools and their own questions, so the low-threshold/high-threshold split this page draws in firms shows up differently: adoption is near-universal (46.6% daily) and the bottleneck still did not move, because what binds is not permission but the next stage of the pipeline.

The gate written as a pass/fail test, and graded against one economy (CIVIC-AI, September 2026)#

When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration — CIVIC-AI 2026 workshop whitepaper, 22 authors including three statisticians from Singapore's Ministry of Manpower, practitioner-opinion with no measurement of its own; the six conditions and the boundary rule are carried on Human-AI Accountability Redesign — restates this page's gate as a definition. AI "genuinely augments work" only when the redesigned workflow clears six conditions, the first of which, durable net value, requires the gain to survive "full accounting of quality, human review, exception handling, rework, recovery, and the cognitive burden shifted to workers," with the failure mode named exactly: "apparent productivity reflects work shifted elsewhere in the workflow." That is the complement constraint converted from an explanation into a test, and its failure mode is the ATLAS bottleneck-migration result above with the sign removed.

The part that adds something this page did not have is §6, the first attempt in the corpus to grade a whole economy against such a test. The statistics are Singapore's, from MOM's Manpower Research and Statistics Department, and they reach the vault secondhand — the whitepaper cites them, does not collect them, and nothing here was checked against the MOM release: 28.5% of firms report having adopted AI; among adopters, 70.7% report improvements in worker productivity; and firms report role redesign (18.9%) more often than reduced headcount (6.2%). The roughly 3:1 redesign-over-reduction ratio is the reorganization-not-substitution claim appearing in a national employer survey, from a labour ministry rather than a payroll vendor — a weaker instrument than Ramp × Revelio's spend-and-records panel (self-report, one small high-capacity economy, no time series) pointing the same way.

The authors' own verdict on their own government's numbers is the more useful half, and it is this page's measurement problem stated by the people who would have to fix it: these figures "provide preliminary evidence relevant to Condition 1 (Durable net value), but do not establish it under our definition, as existing statistics do not capture the full costs of verification, exception handling, rework, recovery, or unofficial AI use." A statistical office with representative employer surveys and linked administrative data cannot currently tell whether its own economy's AI gains are net of the complement costs. Two further readings travel with it. The worker-side conditions fare worse — entry-level PMET openings moved 32,500 (Dec 2025) → 32,800 (Mar 2026) with fresh-graduate outcomes "broadly resilient," which the authors read as not yet showing deterioration rather than as evidence of anything (the same non-finding, on a different instrument, that The Tragedy of the Cognitive Commons now carries from the Canaries revision). And the skills complement shows up as a distribution, not a level: among young Singaporean workers, 38% of those with secondary qualifications used new technology at work against 74% of degree holders, so "human control should therefore be assessed as a capability workers can actually exercise, not only as a formal role assigned on paper."

The complements, measured before the adoption they predict (OpenAI Enterprise, August 2026)#

How Organizations Use AI: Evidence from ChatGPT — Chatterji, Holtz, Rakholia, Tambe & Weeratunga, arXiv 2608.12236, empirical, 69pp — is the first source in this corpus to put a coefficient on the complement claim rather than an analogy or a case. Full treatment at The Enterprise AI Adoption Gradient; what belongs here is the design and the ranking.

The design solves the reverse-causation problem that has dogged every earlier row on this page. Complement stocks are measured in fiscal year 2021, three to four years before the FY2024–25 adoption window, and built as perpetual-inventory accumulations of Compustat flows (SG&A depreciated 20%/yr, R&D 15%/yr) plus the capitalized-software stock, each normalized per employee. A stock laid down in 2021 cannot be an artifact of buying ChatGPT Enterprise in 2025.

FY2021 complement stock per employee → P(adopt)Full sampleExcluding tech / high-R&D
SG&A stock (organizational capital)0.020*** (0.005)0.010 (0.007), n.s.
Capitalized software0.008** (0.003)0.006 (0.006), n.s.
R&D stock0.004*** (0.001)0.003*** (0.001)

(Table 4, reconciled against pdftotext -layout -f 37; the docling parse of that table is shifted and unusable.)

Three things this adds, and one it withholds. It ranks the candidates for the first time, and the winner is the least glamorous one: accumulated SG&A — the crude accounting residue of sales, general and administrative capability — carries five times the R&D coefficient and more than twice the capitalized-software one. It dates them, which no survey instrument on this page can do. And it corroborates the co-invention framing in the authors' own words: "General purpose technologies rarely generate immediate, economy-wide gains; their impact unfolds through a slower process of co-invention in which firms discover use cases, invest in complements, and reorganize production."

What it withholds is the step from adoption to value. These stocks predict who signs the contract, not who gets anything out of it — the paper measures no output, productivity or quality outcome of any kind, and says its estimates "should not be interpreted causally." And the ranking is fragile exactly where generality matters: outside technology and high-R&D sectors, only R&D keeps significance. Add the sample frame — US public companies that bought one vendor's enterprise product, with the control group contaminated by a disclosure-motivated random sampling step that leaves unsampled adopters among the "non-adopters" — and the honest reading is a lower-bound ranking on a selected population, not a decomposition of the complement bundle.

The complement gap also shows up as arithmetic rather than as narrative here, which is new. Conditional on adopting, larger firms use less per employee (WAU per employee −0.032***, output tokens per employee −0.667***) while messages per active user is flat (−0.002, n.s.). The whole size penalty is breadth: a large firm buys the workspace and a smaller share of its workforce ever becomes active. That is "access without redesign" measured directly, on the same firms whose 2021 intangible stocks got them to buy in the first place.

Connections#

  • Autonomous Scientific Discovery — the complements pattern at its starkest: AlphaFold's own creator estimates the technology made structural biology only 5–10% more efficient, because the pipeline's residual stages (toxicity, solubility, experimental validation, clinical testing, regulatory approval) are untouched by better structure prediction

  • The Enterprise AI Adoption Gradient — the firm-level adoption evidence for this page's thesis: FY2021 SG&A, R&D and capitalized-software stocks predicting FY2024–25 ChatGPT Enterprise adoption, and the breadth-not-intensity decomposition of why large adopters use less per head

  • Human-AI Accountability Redesign — where the CIVIC-AI six-condition test is carried in full: the same complement gate written as a pass/fail definition of "augmentation," with the workflow (not the task) as the unit of analysis and a workflow record as the proposed instrument

  • AI Adoption in Scientific Work — this page's thesis measured inside science: task-level time savings that do not become discoveries because the bottleneck moves downstream into physical experimentation, validation and verification of AI output, with the authors citing the same O-ring and task-chaining literature

  • Community Smells Under AI Adoption — the complements thesis measured on team social health rather than output: across 152 professionals the socio-technical benefit of AI in specialization work is entirely mediated by peer consultation, so it is not a property of the tool but of whether its use sustains human knowledge exchange — a complements finding in the strict sense

  • Telemetry vs. Survey Measurement — the instrument caveat that travels with the German HR evidence above, and the one question where the ordering inverts: adoption happening outside any organizational rollout (183 of 410) is invisible to the usage-share and spend-trace instruments the rest of this page relies on, so on shadow adoption the self-report is not the weaker instrument but the only one

  • Controlled Variance: AI's Edge as Reduced Dispersion — the randomized instance of this thesis, and the source of the trust-calibration complement above: automating one production stage improved its output while lengthening the un-automated stage that consumes it, so end-to-end time-to-hire got worse even as offers, starts, and retention improved

  • The Tragedy of the Cognitive Commons — developmental infrastructure as the complement nobody is incentivized to supply: its benefits are non-excludable, so the firm that maintains it is the one at a competitive disadvantage. Also this page's sibling on method: the vault's two no-new-data framework papers, one modelling the substitution decision and one modelling the expertise supply that decision draws on — and they collide on HAT's P7, where upskilling is a private effort choice against a commons the individual does not control

  • The Solo-Authorship Rebound — P1's first labor-composition outcome that breaks sharply at a capability event rather than ramping, in the ledger row above; also the substitution decision observed one level down from the firm, inside a single knowledge-production team, where the replaced party is a coauthor and the "organization" doing the deciding is one researcher

  • Role Averaging, Not Role Elimination — the practitioner counterweight to HAT's extreme-point result: the model's optimum assigns a whole task to one agent, while the observed pattern is roles recomposing around a retained judgment half

  • Task Crossover — the complements argument read off a firm-size gradient: outside-occupation task share falls from 18.9% (2–5 seats) to 16.3% (100+ seats), i.e. AI substitutes least for specialists where the specialists already exist

  • Task Saturation: Broad but Shallow AI Diffusion — the usage-side shape complements predict: AI reaching 68% of occupations but only 21% of their tasks, with ATLAS itself noting that task automation ≠ job automation because "coordination costs, complementary tasks and organizational frictions are highly prevalent"

  • The Household Production Boundary — the limit case of this argument: in the household there is no organization to supply complements and no accounts to measure the output, so 86.5% of AI usage generates value that GDP cannot see by construction

  • AI Product Economics Maturation — the builder-side unit economics of this thesis: ICONIQ prices the complements as internal-AI spend (1–3% → projected 16% of revenue) and names the overrun sources (data infra, enablement, governance) — the intangible-capital complements as budget line items

  • Firm AI-Spend Intensity and Headcount Growth — the firm-level natural experiment: employment gains from AI adoption are intensity-gated (only sustained, material spend beyond chat subscriptions grows headcount) and arrive on a 6–12-month learning curve — this page's complements-lag measured directly on AI-vendor spend + workforce records

  • AI Investment Story, Not Efficiency Story — the startup-financials instance of this thesis: AI companies show lower revenue-per-employee now (gains lag adoption), while AI-native RPE grows faster (complements being built), reconciling the lean-unicorn efficiency claim as a lag

  • Conversation-to-Delegation Shift — the three-population usage gap this page explains; same model, different complements → 99.8% vs 63.3% vs 16.5%

  • Acceleration Whiplash — the downstream-cost evidence of missing complements: when orgs adopt AI faster than they redesign review/QA, throughput rises but quality and incidents degrade — the productivity-paradox failure mode in telemetry

  • Returns to Expertise in Agentic Coding — the cited Hitzig et al. argument that supervision/verification/coordination and domain expertise are the binding complements; the returns-to-expertise data is the worker-level version

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the GDP gradient in reported AI exposure is this page's argument in survey form: lower-income workers may lack the complementary skills/infrastructure (the IMF point) that turn exposure into augmentation rather than replacement

  • The Automation–Optimism Link — optimism and perceived skill-value concentrate among heavy delegators, consistent with complements gating who captures AI's gains

  • Verification as the New Bottleneck — the specific complement that most often binds: review/verification capacity is the redesign agentic AI demands

  • Engineer PM Convergence — the role-redesign complement: jobs shift toward directing, monitoring, and integrating agent output rather than executing tasks

  • Human-AI Accountability Redesign — the accountability/span-of-control complement that has to be rebuilt for delegated agent labor

  • AI Native Product Cadence — the startup-side version: AI-native orgs are born with the complements (workflow, review, tooling) rather than retrofitting them

  • Compounding Data Moat — encoded org-specific context (the systematization complement) is itself an intangible-capital complement in the Brynjolfsson sense

  • Printing Press Software Democratization — capability democratizes broadly, but realized value still concentrates where complements exist; the two together explain the uneven diffusion

  • OpenAI — the lab whose high-complement internal environment is the upper bound of this argument

  • Market-Priced AI Exposure (the AI Premium) — the argument read off equity prices: the AI premium is present in developed markets (17.9 bps/week) but absent in emerging markets incl. China (5.0 bps, insignificant), which Borri-Liu-Tsyvinski attribute to "distance to frontier" — AI risk is systematic, and therefore priced, only where listed firms and investors sit near the complementary AI economy. The complements argument, capitalized

  • AI-Native Organization — the practitioner restatement: Tan's "the leverage is not in the weights, it's in how you wire the work" is this thesis from a stage, and his org-primitive mapping (skills / resolvers / trigger evals) is a concrete enumeration of which complements

  • Balance-of-Power Superintelligence — the same month's other elite statement on whether AI's gains concentrate, from the opposite institutional position (a CEO op-ed vs. 200 academics) and the opposite prescription (distribute the tool vs. rebuild the institutions); neither carries evidence, and the complements argument is the reason to doubt the shared premise that access alone settles distribution

  • Erik Brynjolfsson — the productivity paradox at the top of this page is his, and he organized the open letter that restates it as policy

  • AI and Market Power — the complement measured on personnel records and then written into policy: across Portuguese firms sorted by GenAI exposure, the tertiary-educated share of the workforce runs 0.04 → 0.46 from least to most exposed while productivity nearly doubles (24 → 52 EUR ths/worker), and OECD's conclusion states the complements argument as a competition problem — "if the effective use of AI requires a highly skilled workforce, then barriers to accessing AI-relevant skills may effectively become barriers to entry in AI-intensive markets." Human capital as an entry barrier is this page's thesis with the sign of the harm changed

  • Post-Scarcity Macroeconomics — the complements argument is the standing objection to a quasi-infinite economy arriving on schedule: Musk's forecast prices capability directly into output, with no gate for the workflow, skill and org-design changes this page finds binding at every measured step

  • Pilot-to-Production Gap — the practitioner-side account of why the complement gap stays invisible until it bites. Anthropic × Accenture (vendor-claim) itemize the six conditions that make an enterprise AI pilot succeed — curated data, handpicked AI-native engineers, protected budgets, narrow scope, hidden manual intervention, no downstream stakeholders — and observe that these are precisely the complements production will not supply, so the pilot measures a system nobody will run. That converts this page's thesis into a measurement critique: the pilot is not weak evidence about production, it is evidence about a differently-complemented system. Their survey figures state the same gap as a statistic (64% of organizations past pilots into production across multiple functions, 7% with the data readiness to scale) — one vendor's self-published survey of its own prospective buyers, where this page's core evidence is the Codex three-population natural experiment, so the corroboration is directional only

  • Codified vs Tacit Knowledge Exposure — the complement side of this page's thesis with an outcome attached: in four years of ADP payroll, employment grows faster for experienced workers in tacit-knowledge occupations and falls for young workers where AI usage is automative. It is the closest the corpus comes to measuring which kind of human capital AI complements, and its honest limit is that the substitution half of the index is collinear with years of schooling

Open Questions#

  • The "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign workflows quickly, or does the low cost of tool adoption mask slow, expensive process redesign (the real complement)? Partially answered — and the split is between the two halves of the question. Kalff & Simbeck find both happening at once in the same 410 firms: the low-threshold half diffuses faster than the organization, with 183 of 410 respondents using AI informally on personal devices regardless of employer policy, while the half that needs process redesign stalls exactly where the electrification analogy predicts — advanced analytics "seldom economically or logistically viable" without centralised data and standardised processes, and 20.2% of departments using no AI tool at all. So tool adoption does mask the absence of process redesign, but not by making it look fast: the two run on separate tracks, and the visible one requires no organizational change to happen. Still short of settling it — self-reported, cross-sectional, one function, one country, and no measure of redesign speed where it does occur. #oq/source Extended (2026-09-22), with the first national-survey reading: When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration (practitioner-opinion, citing Singapore MOM statistics it did not collect) reports firms naming role redesign (18.9%) about three times as often as reduced headcount (6.2%) at 28.5% adoption — the redesign half of this question happening at population scale rather than in one favorable case. It does not close the question, and the authors supply the reason: employer-survey statistics "do not capture the full costs of verification, exception handling, rework, recovery, or unofficial AI use," so "role redesign" as a survey checkbox cannot distinguish genuine process redesign from re-labelled tool adoption — which is the exact distinction the bullet turns on. What it does establish is that the distinction is not merely this vault's worry: a labour ministry co-signs the claim that its own instruments cannot resolve it. One more corroboration of the slow half, weak in tier but strong in shape (2026-09-22): Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals treats non-redesign as the field's normal condition rather than a defect — it sets its depth thresholds so that no cohort can read "healthy" at the recurring-use-to-workflow-integration boundary, on the stated belief that the ~2% workflow-embeddedness rate (ActivTrak, secondhand, 120,620 workers) is "an industry-wide cliff, not a cohort-specific defect." practitioner-opinion with no real deployment data, so it is a design assumption and not a measurement; what it adds is that an instrument builder priced the redesign half as near-zero and made that the default. The intensive-margin half, measured rather than surveyed (2026-09-23): How Organizations Use AI: Evidence from ChatGPT (empirical, OpenAI administrative records, 1,764 organizations / 17.4M messages) shows tool adoption and organizational reach coming apart inside the same firms, in the direction this bullet predicts. Among ChatGPT Enterprise adopters, larger firms use less per employee — weekly active users per employee −0.032*** and output tokens per employee −0.667*** on lagged log employment — while messages per weekly active user is flat (−0.002, not significant). The entire size penalty is breadth, not intensity: the workspace is bought, and a smaller share of the workforce ever shows up. That is tool adoption masking the absence of reach, stated as a coefficient. The paper's own conclusion is this bullet's premise in the authors' words — "adoption is only the beginning of deployment" — and its growth decomposition shows the other half moving too: output tokens grew ~7x across all customers between June 2025 and March 2026 and ~4x within the cohort that had already adopted by June 2025, so roughly half of the growth is deepening inside existing adopters. Still short of settling it, and for a reason no sample size fixes: the paper measures no redesign of any kind — no process, workflow, or organizational-change variable exists in it, and no outcome does either — so it can show that reach lags adoption but cannot say whether redesign is what closes the gap.
  • Which complement is the true binding constraint — access/permissions, skills, or review capacity? The paper lists all; it doesn't decompose their relative weight. Sharpened, not answered: the list itself is incomplete. Kalff & Simbeck's German evidence adds an institutional complement (works-council co-determination under BetrVG §87(1) no. 6, EU AI Act high-risk classification) that is not internal to the firm at all and that determines which capability is adoptable rather than how well it is used — so any decomposition needs a fifth term whose weight varies by jurisdiction rather than by firm. Their own most-cited internal blocker is data centralisation, which is closest to "access." A candidate term the list does not have, from a framework that cannot weight it (2026-09-22): When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration adds oversight-competence maintenance — reviewer competence "does not persist automatically. It requires organisations to deliberately assign workers enough substantive review work and enough exposure to AI failure to keep the verification skill current." On that reading "skills" and "review capacity" are not two items on the list but one item measured at two times, and the binding constraint is a stock that depreciates rather than a headcount — which changes what a decomposition would have to estimate (a depreciation rate, not a coefficient). It is practitioner-opinion with no data, so it reorders the candidates rather than weighting them. Its one distributional datum points at access/skills: among young Singaporean workers, 38% of those with secondary qualifications used new technology at work against 74% of degree holders. #oq/source The first measured ranking of candidate complements, on the adoption margin only (2026-09-23): How Organizations Use AI: Evidence from ChatGPT regresses ChatGPT Enterprise adoption on fiscal-year-2021 intangible stocks per employee — three to four years before the FY2024–25 adoption window, so the measure cannot be an artifact of adopting — and the coefficients separate: SG&A stock 0.020*** (0.005), capitalized software 0.008** (0.003), R&D stock 0.004*** (0.001), all reconciled against pdftotext -layout (the docling parse of that table is shifted). Accumulated organizational capital beats the literal software-investment measure by 2.5x and the innovation measure by 5x, which points the decomposition at this bullet's "access" and "skills" terms and away from technical infrastructure. Three discounts keep it partial, and the first is decisive. It ranks predictors of adoption, not of value — the question asks which complement binds the gain, and this paper measures no output, productivity or quality outcome at all. Outside technology and high-R&D sectors only the R&D coefficient keeps significance, so the ordering's generality is unestablished. And the population is US public companies that bought one vendor's product, with the non-adopter group contaminated by a disclosure-motivated random-sampling step. Nothing here touches the institutional fifth term or the oversight-competence depreciation rate; Compustat has no line for either.
  • If complements, not capability, gate value, does model progress have diminishing near-term returns until orgs catch up — and how long is that lag for agentic AI specifically? Bounded, not answered (2026-09-23): How Organizations Use AI: Evidence from ChatGPT supplies the first long usage series from inside enterprise buyers and it runs the wrong way for a simple diminishing-returns story — output tokens grew roughly sevenfold between June 2025 and March 2026, about half of that inside firms that had already adopted, and the acceleration in early 2026 hit every adoption cohort simultaneously, which the authors read as a product- or market-level development rather than the normal post-adoption ramp. A cohort-independent acceleration is evidence against "orgs are the rate limiter and capability waits on them" over a nine-month window. Three reasons it does not close the bullet: tokens are an input measure, not value, and the paper has no outcome variable, so rising consumption is equally consistent with a lag in realized returns; the window is nine months against a question about multi-year catch-up; and the series is overwhelmingly non-agentic by the authors' own Appendix A1, so it says nothing about agentic AI specifically, which is what this bullet asks.
  • HAT's P2 (middle-management vulnerability) is the vault's most-quoted-but-least-tested substitution claim, and nothing here can settle it: Ramp × Revelio resolves seniority only to entry-level / non-entry / manager-plus, which cannot distinguish "middle layers thinned first" from "manager-plus grew more slowly than the bottom." Falsifiable with a level-resolved employment panel (Revelio or matched employer-employee data cut by reporting depth, not seniority band) tracking layer counts before and after intensive AI adoption. The same instrument would settle P4, since it could compare regulated against unregulated industries on the same measure — though P4 now has its first field evidence from German HR (see the ledger row), which confirms the direction on self-report and leaves the panel test outstanding.
  • Corollary 5 claims irreversibility: once an AI agent is strictly cheaper on risk-adjusted grounds, the optimal allocation never reverts, given fixed human costs and negligible switching costs. The vault has no case either way, and this is falsifiable only by a future event — a documented instance of a firm re-staffing with humans a role it had already automated, for reasons other than a regulatory shock or a rise in AI risk (both of which the corollary's own conditions exempt). Watch for it in the same firm-level adoption panels; a reversal with human costs and regulation unchanged would falsify the corollary, and a long clean run of non-reversal would be weak support for it.

Sources#

  • The Shift to Agentic AI: Evidence from Codex — §2 Related literature (organizational complements; David 1990; Brynjolfsson et al. 2019; Hitzig et al. 2026; Demirer et al. 2026a/b); §5 intro (electrification analogy); §6 Conclusion
  • Beyond Benchmarks 2026: Five Data Sets Grounded in the Real World — Emergence Capital, Beyond Benchmarks 2026 (June 2026): the AI-vs-non-AI revenue-per-employee gap as company-financials corroboration of the complements-lag thesis. Note: several tables in this PDF-derived raw have collapsed multi-value cells (the Core Four table stacks Top Decile over Top Quartile into one cell — 736% 184%) — the RPE figures quoted here were re-read from the source PDF page 19 and are those corroborated in the report's prose and slide charts
  • A New Look at AI's Impact on Jobs: Firm-Level AI Spending and Workforce Adjustment — Kharazian, Simon & Stevens (Ramp × Revelio, June 2026): the intensity threshold and 6–12-month learning curve as firm-level, spend-based corroboration — §6 Results, §7 Conclusion (points 4–5)
  • “We Must Act Now”: Sixteen Nobel Laureates Join Leading Economists and AI Researchers in Call to Prepare for AI’s Economic Transformation — Matty Smith, Stanford Digital Economy Lab news release, 2026-07-13 (practitioner-opinion, 558 words, no measurement): Brynjolfsson's "complement humans rather than simply imitate them," Korinek's decades-vs-a-few-years timeline claim, Spence on magnitude/timing uncertainty, Cunningham's "driving in the fog." Carried as a dated marker of elite economic opinion, never as evidence about complements. The statement text and live signatory list are at wemustactnow.ai and are not ingested
  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, Voice AI in Firms (arXiv 2607.28222, 2026-07-30; empirical, pre-registered RCT, 67,056 randomized applicants): §6.2 (recruiters' signal-weighting interaction — the trust-calibration complement), §7.1 (time-to-hire decomposition, 2.62→7.24-day evaluation stage, 20→24 days end-to-end), §8 Conclusion (the complementary-adaptation sentence). Parse warnings and full treatment at Controlled Variance: AI's Edge as Reduced Dispersion.
  • State of AI 2026: The Builder's Economy — ICONIQ Growth, State of AI 2026: The Builder's Economy (2026-07-08, empirical survey with prediction-grade forward figures): §"AI for Internal Productivity" — internal AI spend 1–3%→11%→16% of revenue, the three overrun sources (tokens, data infrastructure, organizational enablement), and the <30% agentic-tool productivity gains that gate the payoff
  • The Human-AI Substitution Principle: When will you be replaced by AI in your organization? — Bonny Banerjee (Independent Scholar) & Shreya Singh (Chennai Mathematical Institute), The Human-AI Substitution Principle: When will you be replaced by AI in your organization?, arXiv 2607.20781 (2026-07-22, 68pp). practitioner-opinion — a formal model with no empirical data, tiered alongside The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise as the vault's other no-new-data framework paper; every result is a theorem conditional on Assumption 1, and the paper says so itself. §3 the HAT model (span-of-control level sizes eq. 1; human cost eqs. 2–4; AI cost eqs. 5–9; the three-component AI risk decomposition eq. 8); §3.4 Assumption 1 (i) Δ$' ≤ Δ$, (ii) T_k/n_k → 0, (iii) r'_0 ≤ r_0, with its own "maintained modelling premise" hedge; §3.7 Table 1 (parameter → observable-construct map, incl. the r'_0 gloss); §4.2.2 Theorem 5 (the Substitution Principle); §4.2.3 Theorem 6 (extreme-point optimality); §4.2.4 Theorem 7 + Corollary 1; §4.3 Theorem 8 + Corollary 2 (middle-management vulnerability, conditional on single-crossing); §4.3.2 Corollary 3 (the μ* protection/vulnerability threshold); §4.4 Theorems 10–12 (risk-adjusted substitution, hybrid optima, full-automation condition); §4.5 Theorems 13–15 + Corollaries 5–6; §5.1 the V = B − C − λR value identity that fixes B; §5.2 accountability mapped onto the three risk components; §5.4 Table 2 (organizational mechanisms); §5.6 Table 3 (predictions P1–P7); §5.7 the calibrated five-level firm and μ* ≈ −1.52; §5.8 the ten stated limitations. Parse notes: Table 1's λ row was truncated mid-cell at ingest (full value: "Firm's tolerance for operational, regulatory, or strategic risk.") and a stray <formula><loc_…> wrapper duplicating an equation block was cleaned during ingest with the math preserved; Tables 2 and 3 were verified verbatim and are the only tables quoted here. The F(i) feasibility values behind the middle-management illustration are from §5.7 prose, and are the paper's own hand-set assumption rather than a derived or measured quantity
  • Return of the solo author: The changing division of labor in science in the age of generative AI — Akira Matsui, arXiv 2607.10780 (2026-07-12; empirical, 300M+ OpenAlex works, 26 fields): §Results (the 23-of-26 positive breaks and the field ordering), SI §L (the balanced-venue attenuation +1.72 → +0.75 pp/yr), SI §M (the event study, −2.9 pp in 2022 then +0.1 in 2023). Cited here only for the P1 ledger row; identification caveats and full treatment at The Solo-Authorship Rebound
  • AI-Augmented Human Resource Management? Insights from German companies — Yannick Kalff & Katharina Simbeck (HTW Berlin), AI-Augmented Human Resource Management? Insights from German companies, arXiv 2607.13839 (v1 2026-07-15 / v2 2026-07-20, 30pp; empirical, mixed methods). §3 research design (14 expert interviews, 3 works-council-adviser group discussions, survey of 427 HR managers fielded February 2025 by an ISO-20252:2019-certified provider, pre-tested on 15 in December 2024, 410 valid, descriptive analysis in R); §4.1 + Figure 1 (tool-use distribution incl. Nichts davon 20.2%, and the 183-of-410 informal-use finding); §4.2 + Figure 2 (the four rationale clusters and their weighted scores; the DEV3 no-personal-data quote; the works-council externalisation channel; the micro-management tension with self-organisation); §4.3 (clerical layoffs at COM1, entry-level career paths, dual competencies); §5 Discussion (the Industry 4.0 / IoT precedent; the stated limitations — cross-sectional, Germany-only, self-reported quantitative measures). Funded by the Hans Böckler Foundation (German trade-union confederation research foundation), grant 2022-797-2; authors declare no conflict of interest. Both figures were opened and reconciled against the prose — Figure 1's bars match the text's "over 20% use no AI tools," and Figure 2's rationalisation dominance matches §4.2; note Figure 1's caption says "Number of mentions per tool" while its axis and labels are percentages, and the percentages are what is quoted here. Parse warnings. (1) This document required a glyph-level repair at ingest: docling rendered every digit and every "DOI"/"URL" label as literal glyph names (/two_os, /D_SC) — 3,944 tokens, hitting N=410, every date, every percentage and all 271 references. The mapping was deterministic and 1:1 and was restored by substitution, then re-verified against the source PDF (N=410, "427 HR managers", "February 2025", and a sampled DOI all confirmed). The staged numbers are correct, but every figure quoted from this source traces through that repair. (2) Table 1 (studies × employee life cycle) is row-shifted: a trailing citation truncates to "Zhai," with its tail leaking onto the next row, so the Development, Management and Offboarding source cells are misattributed — that table is cited nowhere on this page and its content is taken from prose instead. Table 2 (the interview sample) verified clean. (3) Source-internal citation error, flagged rather than repeated: the paper places high-risk HR applications under "Article 35 of the EU AI Act" and cites the European Commission's 2021 proposal rather than the adopted Regulation (EU) 2024/1689. The substance is right — employment and worker-management uses are high-risk under Annex III — but the article number matches neither text
  • Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28, practitioner-opinion): the 30–40%/60% complement arithmetic attributed to Brynjolfsson and McAfee, the business-not-AI KPI answer, and the embedded recruiting-engineer example. No data offered; auto-caption transcript
  • Deploying AI from pilot to production: A practical blueprint for CIOs and technical leaders — Deploying AI from pilot to production, Anthropic × Accenture, 2026-09-11, 38pp, vendor-claim. Cited here for the pilot-insulation account of the complement gap and the 64%-vs-7% data-readiness split (Accenture, AI-Ready Data for Advanced AI, May 2026 — self-published, no methodology). Full treatment and evidence limits on Pilot-to-Production Gap
  • Google AI & Economy ATLAS: AI in Science (September 2026) — AI in Science: Early Insights (Google, Google DeepMind, MIT FutureTech, September 2026, 42pp; empirical with vendor COI). Cited here for §5.2–5.3 and §6.2: the time-saving and reinvestment split, the bottleneck-migration and backlog numbers, the verification tax, and the discussion's explicit O-ring / task-interlock framing. All survey figures are self-reported from a screened non-probability panel; full treatment on AI Adoption in Scientific Work
  • When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration — Wu, Ziems et al. (22 authors incl. Singapore Ministry of Manpower), CIVIC-AI 2026 workshop whitepaper, arXiv 2609.12482, 2026-09-11, 8pp, practitioner-opinion, no original measurement. Cited here for §6 only (the Singapore application: adoption, productivity, role-redesign vs headcount, entry-level PMET openings, and the technology-use gap by qualification) — all third-party MOM/NUS SSR statistics reaching the vault secondhand and not independently verified against the underlying releases; treat every figure in that section as a citation, not a measurement. Tables 1–2 of the paper were reconciled against pdftotext -layout (exact); §6's figures are prose, not table cells. Six-condition framework and boundary rule at Human-AI Accountability Redesign
  • How Organizations Use AI: Evidence from ChatGPT — Chatterji, Holtz, Rakholia, Tambe & Weeratunga (OpenAI + Columbia + Wharton), How Organizations Use AI: Evidence from ChatGPT, arXiv 2608.12236 (2026-08-12, 69pp, empirical). Cited here for §4.2.4 and Table 4 (the FY2021 SG&A / R&D / capitalized-software complement stocks and their adoption coefficients), §4.2.2 and Table 2 (the breadth-not-intensity decomposition of usage by firm size), §4.1 (the 7x / 4x growth decomposition) and §5 (the co-invention conclusion). Tables 2 and 4 were reconciled cell-by-cell against pdftotext -layout pages 35 and 37 — the docling parse of both is row-shifted and the raw carries unrepaired table-collapse / table-weld warnings. Conflict of interest: no independent author — three authors are OpenAI staff and the two academics contributed "in their capacity as paid contractors for OpenAI." Descriptive only, no outcome measured, sample is ChatGPT Enterprise buyers. Full treatment, the two source-internal prose↔figure disagreements, and the sample-selection notes at The Enterprise AI Adoption Gradient
§ end
Cited by 43
Related articles