H
Howardism
Plate IIAI Economics & LaborHOWARDISM

Controlled Variance: AI's Edge as Reduced Dispersion

PublishedAugust 4, 2026FiledConceptDomainAI Economics & LaborTagsWorkforceEconomicsHuman AI CollaborationMeasurementEmpiricalReading28 minSourceAI-synthesised

Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants between AI voice interviewers and human recruiters — offer rate 8.70%→9.73% (+12%), job starts +18%, one-month retention +18%, no productivity decline, with humans making every hiring decision in both arms. The mechanism the authors name is *controlled variance*: the AI follows the firm's interview protocol more consistently (topic order τ 0.53 vs 0.33, question similarity 0.59 vs 0.43, significantly lower cross-interview variance) while still adapting per applicant and using *richer* vocabulary — AI wins by being less dispersed, not more capable. The wiki's only randomized causal estimate of AI substituting for a human in an expert conversational task

Illustration for Controlled Variance: AI's Edge as Reduced Dispersion

Sources#

Summary#

Controlled variance is Brian Jabarian and Luca Henkel's name for the mechanism by which an AI agent beat human experts at a human-intensive task: not by being smarter at any single interview, but by executing the same protocol with less dispersion across interviews while remaining responsive within each one. The claim is that AI's advantage in information collection is a variance property, not a mean-capability property — "AI can improve firm outcomes by reducing variance in information collection."

The evidence is a pre-registered natural field experiment (AEA RCT Registry #15385, UChicago IRB) run with PSG Global Solutions, a recruitment-process-outsourcing subsidiary of Teleperformance, over March 7 – June 7 2025: 70,884 applications for entry-level customer-service jobs in the Philippines, 67,056 randomized into three arms — AI Interviewer (40,103), Human Interviewer (13,557), Choice of Interviewer (13,396). In every arm a human recruiter reviews the transcript, audio, and test scores and makes the hiring decision. Only who conducts the interview varies. That design is the whole point: it cleanly separates the information-collection stage (automated) from the evaluation stage (not automated), and measures what happens to firm outcomes when only the first is handed over.

Evidence note. empirical, and the strongest causal design in this vault's economics cluster: randomized assignment, pre-registered, administrative outcome data (offers, starts, retention, separations, performance scores) linked to interview transcripts — not telemetry, not a survey, not a projection. Two things to hold alongside it. (1) Disclosure: the pre-registered data collection finished 2025-06-07; two months later, on 2025-08-01, Jabarian accepted an unpaid Chief Economist role at the partner firm, renewing the research partnership for five years under the original publication rights. The partner supplied data but had no role in analysis or the publication decision. The role postdates the data; it does not postdate the writing. (2) The AI system is never identified — no model, no vendor, no version. Google Cloud supplied infrastructure and engineering support. So the result is not attributable to, or reproducible against, any named model.

What the experiment establishes#

All figures below are from the paper's prose (see the parse note in Sources). The unconditional sample compares everyone randomized into AI Interviewer vs Human Interviewer — an intention-to-treat contrast, so it is net of the AI's own failures.

Outcome (all randomized applicants)HumanAIEffect
Received a job offer8.70% (1,179/13,557)9.73% (3,904/40,103)+1.03pp, +12%, p<0.001
Started the job5.65%6.71%+18%, p<0.001
Employed ≥1 month4.97%5.85%+0.88pp, +18%, p<0.001
Employed ≥2 months+17%, p<0.001
Employed ≥3 months+16%, p=0.003
Employed ≥4 months+17%, p=0.015

Conditioning on applicants who accepted an offer — which removes the offer-rate channel and asks whether the marginal AI-sourced hire is worse — the gains shrink but survive at short horizons: job starts 68.84% → 73.36% (p=0.003), employed ≥1 month 60.60% → 64.32% (p=0.025), ≥2 months 55.62% → 59.13% (p=0.038). At three and four months the conditional differences are still positive (+5%, +7%) but no longer significant on a shrinking sample (p=0.16, p=0.25). So the effect is not simply "more offers to worse candidates."

Three checks that matter more than the headline:

  • Robustness. Adding gender, application source, pre-treatment engagement score, repeat-applicant status, plus week / recruiter / city / job-posting fixed effects barely moves the estimates (offer effect 0.0104 → 0.0101). Clustering standard errors at the recruiter rather than applicant level widens them but keeps offers, starts, and one-to-three-month retention significant; the four-month effect loses significance under recruiter clustering (0.0038, ns). The long-horizon tail is the weakest part of the result.
  • Not a few outlier recruiters. Recruiter-level offer rates correlate ρ=0.87 across arms, and 69% of recruiters extend offers at a higher rate for AI-interviewed applicants. The effect is broad, not driven by a handful of AI enthusiasts.
  • Achieved despite a 12% failure rate inside the treatment arm. 7% of AI interviews aborted on a technical failure of the voice agent, and 5% of applicants ended the call because they were unwilling to speak to an AI. Those failures are inside the AI arm and drag the ITT estimate down. The +12% is net of them.

What it does not establish#

Precision here is the difference between reading this paper correctly and over-reading it.

  • An offer-rate increase is not a hiring-quality increase. The paper's quality proxy is retention ≥1 month, and it is candid that this is a proxy — justified by low frictions on both sides in a high-turnover BPO market (up to ~60% annual attrition), where firms fire poor performers quickly and workers quit without penalty. It is also, not incidentally, the metric the firm is paid on: clients compensate PSG for hires who stay at least a month. The outcome that improved is the recruiter's own revenue metric.
  • Productivity is a null, not a gain, and on a subsample. For workers where it is observed, there is no significant difference in average handle time, employer quality-assurance score, or customer-satisfaction score. The paper's own reading is the right one: the effects are "not accompanied by decreased employee productivity." That is an absence of harm, not a demonstrated quality improvement. Separation composition is likewise unchanged (voluntary 58.25% human vs 58.81% AI, p=0.95).
  • Nothing beyond four months is measured, and the four-month estimate is the one that fails recruiter clustering. No promotion, wage-growth, client-value, or long-run performance outcome exists.
  • The mechanism is descriptive, not causal. Randomization identifies the treatment effect on offers, starts, and retention. The transcript analysis — topic coverage, order adherence, applicant linguistic features — is observational description of what differed between arms, and the paper consistently says "associated with." The controlled-variance story is the most plausible account of why; it is not itself identified.

The mechanism: consistency without mechanicalness#

The firm gives recruiters structured guidelines — 14 possible topics, a recommended order, sample questions — with substantial latitude to tailor. The AI agent gets the same guidelines. From 34,109 transcripts:

Interviewer behaviorHumanAI
Share of the 14 topics covered38%45% (p<0.001), and lower variance (Levene p<0.001)
Adherence to guideline topic order (Kendall τ)0.330.53 (p<0.001), lower variance
Similarity of questions asked to guideline questions0.430.59 (p<0.001)
Interviewer vocabulary richness6.667.64 (p<0.001), tighter distribution

Two features of this table carry the concept. First, neither party mechanically reads the script — τ=0.53 and similarity 0.59 mean the AI deviates constantly, tailoring follow-ups to each applicant. Consistency is across interviews, not within them. Second, the consistency does not come at the price of flatter language: the AI's vocabulary richness is higher than the humans'. The usual trade-off — you can have standardization or responsiveness — does not appear.

The variance claim needs one qualification the figures make plain. Figure 2's distributions show the AI's topic coverage is not tightly clustered — it is sharply bimodal: roughly 20% of AI interviews sit in the lowest bin (0–10% of topics) and ~42% pile into a single high bin (70–80%), with little mass between. The human distribution is broader but smoother. The same shape appears in topic-order adherence (a ~46% spike at τ 0.75–1.00 alongside a ~26% mass near zero). So the AI's consistency is conditional: given the call functions, it executes the protocol very uniformly; when it fails — the 7% technical aborts, 5% AI refusals, plus unavailable and disengaged applicants — it fails at near-zero coverage rather than degrading gracefully. "Lower variance" is measured across a mixture of a tight high mode and a discrete failure mode, which is a different object from "more reliable everywhere." It is arguably the better shape for a firm (failures are legible and cheap to route to a human), but it is not the smooth variance reduction the phrase suggests.

Benchmarked against the distribution of individual human recruiters, the AI agent beats 100% of them on question-similarity to guidelines, 83% on topic-order adherence, and a more modest 61% / 64% on topic coverage and vocabulary richness. It is not superhuman at any of these; it is reliably above the middle at all of them at once.

The proposed link to offers runs through applicant language in a two-step design: (1) within the human arm, regress offers on eight standardized linguistic features — number of exchanges, applicant vocabulary richness, and syntactic complexity predict offers positively; backchannel cues and applicant-posed questions predict negatively; (2) compare features across arms — the positive predictors are higher in the AI arm and the negative ones higher in the human arm. More structure ⇒ more of the signal recruiters reward.

The margin a substitution model cannot price#

Worth stating flatly, because the formal literature on when AI replaces a worker is built on the other margin. Banerjee & Singh's HAT model (arXiv 2607.20781, practitioner-opinion, no data) reduces every substitution decision to a comparison of risk-adjusted scalars — C'_ik + λR'_ik < C_ij + λR_ij — and that condition has no way to express a variance advantage:

  • Output quality is assumed identical. The objective is derived from V = B − C − λR with B dropped because "B is fixed for a given task." Both agents deliver the same benefit by construction; only cost and risk differ. Controlled variance is a claim about the dispersion of output quality across repetitions, and the model has neither repetitions nor variable quality.
  • R is a level, not a second moment. C + λR looks like a mean-variance functional, but R is defined as error/turnover/absenteeism (human) and reliability/compliance/reputational failure (AI) — an expected-loss level. No variance term appears in the paper.
  • A generous encoding collapses the distinction anyway. Setting R' < R would say the AI is more reliable — but that is a mean shift in effective cost, indistinguishable from the AI being cheaper. The model cannot separate "better on average" from "less dispersed," which is exactly what this experiment separates: the AI beat 61%/64% of individual recruiters on topic coverage and vocabulary richness while beating 100% on question-guideline similarity and 83% on order adherence.

The direction is the sharper problem. Every worked example in that paper runs R' well above R (calibrated $20k vs $5k), and all seven of its predictions turn on λ(R' − R) > 0 — AI as the riskier party whose risk penalty is what delays substitution. Here the AI's reliability was the product, not the obstacle. A second structural result fails on the same evidence: that model's optimum is an extreme point (the whole task to one agent, mixtures only from cost ties or externally imposed constraints), while the arrangement that actually produced +12% offers was a split of one task by stage — AI collects, human decides — which it can represent only as a constraint imposed from outside the optimization.

The channel the paper does not isolate: discretion to abort#

Classifying transcripts into ten interview types surfaces one gap much larger than any linguistic difference: 25% of human interviews are "screen-outs" — ended by the recruiter because the applicant doesn't meet a requirement — against 7% of AI interviews. (Comprehensive interviews are near-parity: 39% human vs 42% AI.)

The paper folds this into the standardization story. It is worth naming separately, because it is a different mechanism: a human recruiter exercising discretion to stop removes an applicant from the funnel before the test and evaluation stages, mechanically suppressing offers; the AI, following protocol, largely does not stop. On this reading part of the +12% is "the AI does not quit," which is a discretion difference rather than an information-quality difference. It is still variance reduction — the recruiter's stopping rule is exactly the kind of high-discretion behavior that varies across people — but "we removed the interviewer's option to give up early" is a considerably more portable and less flattering description of the intervention than "AI collects better information." The paper reports the number and never decomposes the two channels. See the open question below.

Recruiters score AI higher and trust it less#

The human evaluators' response is the paper's most interesting behavioral finding, and it points two ways at once:

  • Higher scores. Mean interview score 1.90 (human-conducted) → 2.01 (AI-conducted), p<0.001 — driven entirely by a shift from score 1 to score 2, with the share of 3s unchanged. AI lifts the floor, not the ceiling. Sentiment of recruiters' free-text justifications shifts the same way (positive 24% → 31%; negative 38% → 28%).
  • Lower weight. Regressing offers on standardized interview, language-test, and analytical-test scores, the interview score is significantly less predictive of offers in the AI arm (interaction −0.047, then −0.029 with controls, both significant) while the language-test score is more predictive (+0.028 / +0.022). And the discount is concentrated exactly where the theory says it should be: it is significant only among the recruiters who told the survey that interview performance matters more than test scores.

So recruiters simultaneously rate AI-collected interviews more favorably and lean harder on the independent signal when deciding. They discount the AI's signal — and the paper's conclusion is that this is a complement not yet built: "realizing the full returns from automation requires complementary adaptation in how humans rely on AI signals for decision-making."

Their forecasts were also wrong in a specific direction. Before results were disclosed, 36% of recruiters expected AI-interviewed applicants to receive lower offer rates (49% equal, 15% higher), 48% expected lower retention, and 61% expected AI-led interviews to be of lower quality. Only 12% of recruiters expected AI's impact on themselves to be generally positive — against 47% of the applicants. Incumbent expertise predicted the sign wrong.

Automating the interview lengthens the hire#

Decomposing time-to-hire (Remote mode, successful hires only) produces the cleanest causally-identified instance of bottleneck relocation in the vault:

StageHuman (median)AI (median)
Expression of interest → interview0.51 days0.32 (p<0.001)
Interview → offer decision2.62 days7.24 (p=0.005)
Offer → job start12.00 days13.00 (p=0.55, ns)
End-to-end20.00 days24.00 (p=0.033)

Automating the interview made scheduling nearly instant (the agent is available 24/7) and made the human evaluation step ~2.8× slower, because recruiters now review a conversation they did not have. Net, the process got longer. The firm bought a better-quality funnel and paid for it in latency at the one stage it did not automate — Verification as the New Bottleneck observed under randomization rather than inferred from telemetry.

Applicants: no backlash, less naturalness, and a perception/outcome split on discrimination#

  • Offer acceptance is statistically unchanged: 93.64% human vs 92.14% AI (p=0.086 — directionally lower, not significant at 5%). NPS 8.84 vs 8.97 (p=0.25).
  • Perceived quality is a wash except on naturalness. No differences in perceived stress, comfort, follow-up flow, feedback frequency, recruiter knowledge, or fairness. AI questions are rated slightly more relevant (p=0.044). AI interactions are rated significantly less natural (p=0.014), which pulls the composite interview-quality index in the humans' favor (p=0.076).
  • Reported gender discrimination nearly halves: 5.98% (human) → 3.30% (AI), p=0.020. But the realized offer gap does not move. Female applicants receive 20% fewer offers in the human arm (7.94% vs 9.86%) and 17% fewer in the AI arm (8.98% vs 10.87%) — and both gaps vanish once controls, fixed effects, and test scores are added, implying they are composition rather than recruiter discrimination. The honest summary: AI halved the feeling of being discriminated against without changing the measured gap, which cuts against both the "AI will amplify bias" fear and any claim that it fixed anything. Small samples on the perception item (62/1,818 and 22/346).
  • Open-ended sentiment favors AI (71% vs 52% positive, p=0.005) but only 9.6% responded, and the paper flags the selection risk itself.

Choice, and negative sorting#

Given the choice, 78.4% of applicants pick the AI (69.3% walk-in, 81.6% remote), stably over the whole experiment. Attitudes predict it: among survey respondents expecting AI's workplace impact to be positive, 77% choose AI; balanced, 72%; negative, 65%. Women choose AI slightly more than men (80% vs 76%, p<0.001).

But choice sorts negatively on measured ability: applicants who choose the AI score lower on both the language test (3.14 vs 3.37 out of 6) and the analytical test (47.54 vs 49.77 out of 100), both significant. Comparing choosers to those assigned the same interviewer weakens this (three of four differences lose significance with controls), so the sorting is real but modest. The practical implication for a firm offering the choice: the convenience that makes AI popular is most attractive to the applicants the firm is least eager to hire.

Where the economic case actually rests#

The paper's cost appendix benchmarks nine wage × AI-price settings. AI interviewing is cost-effective in 8 of 9 statically, breaking even after ~8,500–24,000 interviews in low-wage environments, ~4,500–7,000 in mid-wage, and ~2,000–3,000 in high-wage; adding a 5% failure rate raises break-even counts 7–57%, worst where wages are moderate and AI prices high. So on cost alone, "AI adoption scales most rapidly where human labor is expensive."

The experiment itself ran in the cheapest labor environment of the three — Philippine customer-service roles paying ₱16,000–25,000/month (≈$280–435). That is the setting where the pure labor-cost argument for automation is weakest. Whatever this firm gained, it did not primarily gain it by undercutting a wage.

Boundary conditions the authors state: benefits should concentrate in "high-volume, high-turnover environments where tasks are repetitive, outcomes are rapidly observable, and, importantly, where variance in human performance imposes costs" — and settings depending on "tacit knowledge, relational inference, or screening of highly specialized skills may benefit more from human screening." Every one of those conditions holds in this setting and most of them fail in senior or specialized hiring. The result is about entry-level, high-volume, quickly-observable screening; it should not be read as a claim about interviewing in general.

The same task, observed instead of randomized#

Kalff & Simbeck's N=410 survey of German HR departments (arXiv 2607.13839, empirical) covers the same functional domain with the opposite instrument, and the two do not so much disagree as answer different questions. This experiment asks does it work, conditional on deployment, at a site chosen because deployment was possible. The survey asks do firms deploy, across a population that mostly has not: in Germany applicant pre-selection (Bewerber:innen-Vorauswahl) is the second-most-adopted HR AI tool at 35.6%, but the specific intervention randomized here — an AI conducting the conversation (Virtuelle Gespräche) — sits at 17.1%, with candidate matching 14.9% and interview-question generation 10.5%.

The stated reason for declining is a variable this design holds fixed by site selection: the labour market runs the other way. German firms report reluctance to invest in AI ranking or filtering because "complex AI-driven screening systems may inadvertently exclude suitable candidates in sectors already struggling to find qualified applicants" — they are "concerned about missing out on scarce talent or see limited economic justification for such investments." That is a cost asymmetry between error types: where applicants are scarce, the expensive mistake is the false rejection, not the bad hire.

It maps onto this paper's own boundary conditions almost term by term. The authors expect benefits where "tasks are repetitive, outcomes are rapidly observable, and, importantly, where variance in human performance imposes costs" — an entry-level Philippine BPO with up to ~60% annual attrition, 70,884 applications in three months, and retention observable in weeks. Invert each condition and the German picture appears: fewer applicants per role, slower-to-observe outcomes, and the cost of dispersion borne mainly as candidates lost. Controlled variance is worth most when what you are tightening is dispersion around a cut you can afford to make. Germany also adds a constraint the experimental site did not face — an AI conducting selection interviews sits in the EU AI Act's high-risk tier, and any system evaluating existing employees triggers works-council co-determination under BetrVG §87(1) no. 6.

Two limits on the comparison: the survey is self-reported and cross-sectional, and its adoption figures describe German HR in general rather than the high-volume entry-level segment where this result claims to apply — Germany may simply have fewer such roles rather than declining the tool for them. The point is not that the experiment fails outside its site. It is that the result and the adoption rate are governed by different variables, and only the second is the one a firm actually decides on.

Connections#

  • Telemetry vs. Survey Measurement — the third instrument, and the only one that identifies causation. This paper runs all three inside one design and they disagree: administrative records say AI-interviewed applicants get more offers and stay longer; the applicant survey says AI interviews feel less natural and scores their composite quality lower (p=0.076); and the recruiter survey's forecast (61% expected worse interviews) is flatly contradicted by the experiment it forecasts. Randomization is what adjudicates — a rank above the telemetry-vs-survey question that page arbitrates
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — a direct experimental test of the interpersonal-bottleneck assumption baked into Frey & Osborne 2017 and Brynjolfsson 2018. Interviewing is a relational, conversational, judgment-laden task, and an AI voice agent conducting it produced better firm outcomes than trained human recruiters. Where Steele & Cruz's query-built instrument already put Social jobs top on usage counts, this supplies outcome-measured causal evidence on one such task — with the paper's own boundary condition (tacit/relational/specialized screening still favors humans) preserving the older intuition where it survives
  • Organizational Complements to AI — the paper's own closing argument, and an unusually clean instance: capability was sufficient to improve the interview, and insufficient to improve the process. Recruiters discount the AI's signal, and the human evaluation queue lengthened 2.62 → 7.24 days, so end-to-end time-to-hire got worse. The missing complement is named precisely — "complementary adaptation in how humans rely on AI signals for decision-making". Also the vault's home for the HAT substitution model this page's mechanism falls outside: the 2.62 → 7.24-day evaluation stage is one measured instance of that model's coordination parameter r'_0 exceeding its human counterpart, against an assumption it needs to run the other way
  • Human-AI Accountability Redesign — the design is the decision-rights split the five-pillar prescription asks for: AI collects, humans decide, and the agent discloses its artificial identity and states that a human will make the call. This is what that page recommends, implemented at 70K scale — and the result shows the split works on outcomes while surfacing the cost, which is that the human decision stage becomes the queue
  • Configurable Human Participation — the field counterpart to HAS-Bench's controlled sweep. HAS-Bench varies human participation on an LLM-simulated human across 397 tasks; this fixes participation at one configuration (agent collects, human holds all authority over the decision) and randomizes 67,056 real people through it. Both find the configuration is what matters, from opposite ends of the realism/control trade
  • The Solo-Authorship Rebound — the contrast on what happens to the residual. Here the retained judgment stage expanded 2.8× (2.62 → 7.24 days); there, a researcher absorbing their coauthors' execution work ends up covering 23% less content ground (breadth 0.089 vs 0.115) with no movement into new territory. Substitution can make the human's remaining half bigger or smaller, and neither case predicts the other — though this one is randomized and that one is an interrupted time series with no untreated unit
  • Role Averaging, Not Role Elimination — a measured instance of the averaging: the recruiter's job did not disappear, it lost the interviewing half and kept the evaluating half, which then grew from 2.62 to 7.24 days of median work per hire. The paper's own framing is "redirecting recruiter expertise toward evaluation" — role recomposition, with the human keeping the judgment-bearing part
  • AI Employee Framing — the mirror image, and it points the other way. Here the AI is framed unambiguously as a tool (identity disclosed at call start, human decision-maker stated explicitly) and the humans respond by increasing scrutiny of the independent signal — the opposite of the accountability diffusion the employee framing produces. Suggestive, not a test: framing was not manipulated
  • The Automation–Optimism Link — the belief→delegation link outside the AEI's tech-skewed sample: in a Filipino entry-level customer-service population (60% female), 47% expect AI's workplace impact on them to be positive vs 19% negative, and that belief predicts choosing the AI interviewer (77% / 72% / 65% across positive / balanced / negative). Recruiters — the incumbents — are far more pessimistic (12% positive)
  • Returns to Expertise in Agentic Coding — a rare case running against the expertise gradient: 131 experienced recruiters forecast the direction of the effect wrong (36% expected lower offer rates, 48% lower retention, 61% lower interview quality), and the task their expertise was built on is the one that automated cleanly, while the task expertise transferred to — evaluation — is the one that stayed human and became the bottleneck

Into hubs (one-way)#

  • Verification as the New Bottleneck — the bottleneck relocation measured under randomization: automating information collection cut scheduling time to 0.32 days and inflated human evaluation from 2.62 to 7.24 days, making the end-to-end process significantly longer (20 → 24 days, p=0.033)

Open Questions#

  • How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them.
  • The AI system is never identified — no model, vendor, or version, only that Google Cloud supplied infrastructure. Is "controlled variance" a property of a 2025-generation voice agent under this firm's prompt, or of AI-conducted interviews generally? Nothing in the paper lets a replication know what it is replicating.
  • Retention ≥1 month is both the quality proxy and the metric the recruiting firm is paid on by its clients. Does the AI advantage survive on an outcome the intermediary is not compensated for — client-side performance at 12 months, promotion, or wage growth? The four-month estimate, the longest horizon measured, already fails significance under recruiter clustering.

Sources#

  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Brian Jabarian (Chicago Booth) & Luca Henkel (Erasmus Rotterdam), Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews, arXiv 2607.28222 (2026-07-30; 119pp, 27 tables; empirical). §2 design and sample (70,884 applications, 67,056 randomized, three arms, human evaluation throughout); §3.1 offer/start/retention effects and the conditional sample; §3.2 separation reasons and the three productivity nulls; §4.1 interview-type classification (25% vs 7% screen-outs; 7% technical failure, 5% AI refusal); §4.2 recruiter-behavior measures and the per-recruiter benchmark; §4.3 the two-step linguistic-feature design; §5 applicant acceptance, NPS, perceived quality, discrimination reports, choice and negative sorting; §5.4 gender heterogeneity; §6 interview scores, sentiment, and the signal-weighting interaction; §7 time-to-hire decomposition and the nine-setting cost benchmark; §8 conclusion (controlled variance, boundary conditions). Figures viewed and reconciled: Figure 2 (six-panel topic-coverage / topic-order / question-similarity distributions by arm — confirms the prose means and reveals the AI's bimodality, noted above) and Figure 5 (cumulative average time-to-hire, whose three points reproduce the prose averages exactly: AI 0.58 / 14.50 / 29.19 days, human 1.86 / 9.28 / 23.14). Per vault convention no number on this page is read off a chart; every figure comes from prose. Pre-registered: AEA RCT Registry #15385; UChicago IRB #IRB24-1894 / #IRB25-1002. Disclosure: the corresponding author accepted an unpaid Chief Economist role at the partner firm two months after data collection closed; the partner had no role in analysis or the publication decision.
  • The Human-AI Substitution Principle: When will you be replaced by AI in your organization? — Banerjee & Singh, arXiv 2607.20781 (2026-07-22; practitioner-opinion, formal model, no empirical data). Cited above only for what its cost conditions can and cannot express: §5.1 (the V = B − C − λR identity that fixes B per task), §3.2–3.3 + Table 1 (R defined as an expected-loss level for both agent types), §4.2.2 Theorem 5, §4.2.3 Theorem 6 (extreme-point optimality), §5.7 (the calibrated R' = $20k vs R = $5k). Full treatment and the P1–P7 ledger at Organizational Complements to AI
  • AI-Augmented Human Resource Management? Insights from German companies — Kalff & Simbeck, arXiv 2607.13839 (2026-07-15 / v2 07-20; empirical, N=410 German HR managers plus 14 expert interviews). Cited above for the adoption side of the same functional domain: §4.1 + Figure 1 (tool-use shares — applicant pre-selection 35.6%, virtual interviews 17.1%, candidate matching 14.9%, interview-question generation 10.5%), §4.2 (the skill-shortage reluctance to invest in AI ranking/filtering; the AI Act and co-determination constraints). Glyph-level repair at ingest (docling emitted all digits and DOI/URL labels as glyph names; restored 1:1 and re-verified against the PDF), so every number traces through that repair; Table 1 is row-shifted and cited nowhere, Table 2 verified clean. Full treatment at Organizational Complements to AI
  • Parse warnings — the treatment-effect coefficients are intact, the bookkeeping rows are not. Every point estimate checked against the PDF-derived tables is byte-exact, and every figure quoted on this page comes from the paper's prose. Three tables in the raw are damaged and should not be read directly: Table 1 (job performance) has its three dependent-variable group headers collapsed into one repeated header, so the table cannot be used to say which column is which DV — the productivity nulls above are taken from §3.2 prose and the table's own Notes. Table B.5 Panel B has its standard errors stripped from the AI Interviewer row and prepended onto the row below; corrected AI Interviewer values are 0.0078***(0.0021), 0.0074***(0.0021), 0.0074**(0.0031), 0.0061***(0.0020), 0.0058***(0.0019), 0.0058**(0.0029), 0.0037**(0.0015), 0.0038***(0.0015), 0.0038(0.0024) — the last column is where the four-month effect loses significance under recruiter clustering. Table B.6 Panel A has its Controls / Clustering / Observations / R² rows scrambled (correct: Controls -/Yes/Yes/-/Yes/Yes; Clustering App./App./Rec./App./App./Rec.; Observations 4,708/4,575/4,575/4,708/4,575/4,575; R² 0.0018/0.0700/0.0700/0.0011/0.0611/0.0611) and Panel B's standard-error row is dropped entirely. Table B.5 Panel A and Table 2 verified clean.
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 13
  • Organizational Complements to AI×5

    Every argument above is inferred — from an adoption gap, from financial ratios, from a spending threshold, from budget-variance anecdotes. Jabarian & Henkel's…

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated×4

    Controlled Variance — the outcome-measured counterpart to every instrument on this page: instead of inferring how exposed a relational task is, it randomizes…

  • Role Averaging, Not Role Elimination×4

    ICONIQ surveys intent and Ambrosino describes his own org. Jabarian & Henkel's hiring field experiment is the same split observed under randomization, on one…

  • AI Employee Framing×3

    The tool-framed version of "Scout," measured. The HBR paper's own example of a real AI employee — "Scout," on a company's HR org chart, "autonomously reviewing…

  • The Automation–Optimism Link×3

    Controlled Variance — the belief→delegation link outside this survey's tech-skewed sample (Filipino entry-level customer-service applicants, 60% female: 77% /…

  • Configurable Human Participation×3

    HAS-Bench sweeps many configurations against a simulated human on 397 tasks. Jabarian & Henkel's hiring experiment is the opposite trade: one configuration,…

  • Human-AI Accountability Redesign×3

    voice ai firms job interviews — Jabarian & Henkel, Voice AI in Firms (arXiv 2607.28222, 2026-07-30; empirical, pre-registered RCT): §2.4 (the AI discloses its…

  • Telemetry vs. Survey Measurement×3

    Controlled Variance — the third instrument above: the vault's only randomized causal estimate of AI substituting for a human in an expert conversational task,…

  • Erik Brynjolfsson×2

    2. Brynjolfsson, Mitchell & Rock (2018) — the SML rubric. Suitability for Machine Learning: crowd-sourced ratings of detailed work activities against a…

  • Open Questions Backlog×2

    Controlled Variance: How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human…

  • Returns to Expertise in Agentic Coding×2

    Controlled Variance — the counter-case, and it cuts two ways. In a randomized hiring experiment, 131 experienced recruiters forecast the effect's direction…

  • The Solo-Authorship Rebound×2

    Controlled Variance — the vault's randomized substitution case, and the contrast on what happens to the residual: there the retained judgment stage expanded…

  • AI Economics & Labor

    Controlled Variance — Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants between AI voice…

Related articles