Sources#
- AI-Moderated Interviews for Market Research and Digital Twins Calibration
- Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
The question#
From Controlled Variance: AI's Edge as Reduced Dispersion (#oq/now, partially answered 2026-08-17 by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators): how much of the +12% is controlled variance in information collection versus removal of the interviewer's discretion to abort? This page also covers the sibling #oq/source question on the same page: the AI system is never identified, so is "controlled variance" a property of this one agent or of AI-conducted interviews generally?
The new input is AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins (Deng, Liu, Toubia & Jain, arXiv 2609.29143, empirical, preregistered). It randomizes an adaptive GPT-5.1 voice moderator (N=139) against a static arm with the same fixed questions (N=154), on the same platform and in the same recruitment batch. A human-moderator arm (N=24) was recruited separately and later.
Short answer#
The question has three parts, not two. Jabarian & Henkel's AI arm differs from human recruiters in three ways at once. It follows the guide more uniformly (standardization). It still tailors follow-ups to each applicant (adaptivity). And it screens out far less often (abort: 7% vs 25%). The 2026-08-17 synthesis bounded the abort part. The new study is randomized evidence on the standardization-vs-adaptivity split, from a different task.
- Information side: adaptivity carries it and standardization alone does not. Deng et al.'s static arm is pure standardization: identical questions, no probing, no interviewer discretion, no aborts. It is the weakest arm on every content measure: 5.01 vs 7.48 needs per interview, and lower breadth in three of four blocks, which survives length control (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins). It is also the more dispersed arm in outcome, as shown below. So the paper's own description of the mechanism, "consistent protocol, tailored follow-ups", is doing the work through its second half.
- Abort side: unchanged. Deng et al. settle eligibility with screeners before the interview, neither moderator could end an interview on a disqualifier, and the paper records no aborts. No other vault source measures a screen-out channel. The bounds on What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators remain the best available: the arms lose about the same share of interviews (43% vs 45%), the symmetric completion-conditional correction gives +16%, and the break-even is r* = 5.8%. The decomposition itself still has no direct measurement.
- Sibling question: weak movement on generality, none on identification. A second, named system (GPT-5.1, temperature 0, on an unnamed startup's platform) shows lower coverage dispersion than expert human moderators in a second task. That is a small amount of evidence that the signature is not idiosyncratic to Jabarian & Henkel's agent. It is on the non-randomized contrast with n=24, and it does nothing to identify the hiring study's system.
1. What the randomized split shows#
The static arm is least discretionary and most dispersed#
Controlled Variance: AI's Edge as Reduced Dispersion measures the hiring AI's advantage in procedural terms: topic-order τ 0.53 vs 0.33, question-to-guideline similarity 0.59 vs 0.43, and lower cross-interview variance on both. A reading of "controlled variance" that takes it literally would predict that the most uniform procedure yields the most uniform information. Deng et al.'s static arm is the limiting case. Every participant gets exactly the same questions, so procedural dispersion is zero by construction.
Breadth dispersion by arm, from Table 7 of the raw (mean breadth with SD in parentheses; the table was verified cell by cell against pdftotext at compile time, see AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins Sources). The coefficient of variation (CV) is this page's arithmetic.
| Block (max breadth) | AI-moderated | Static | AI CV | Static CV |
|---|---|---|---|---|
| Purchase journey (9) | 7.29 (0.91) | 5.90 (1.44) | 0.12 | 0.24 |
| Brand awareness (4) | 3.32 (0.80) | 1.73 (0.53) | 0.24 | 0.31 |
| Brand perception (6) | 5.88 (0.35) | 5.52 (0.69) | 0.06 | 0.13 |
| Category usage (3, AI capped at 3 questions) | 2.94 (0.25) | 2.94 (0.24) | 0.09 | 0.08 |
In the three blocks where the AI could probe freely, the adaptive arm has the lower relative dispersion in all three and the lower raw SD in two. In the block where probing was capped, the arms are identical. So removing all interviewer discretion did not produce uniform information. The respondents' own variation passes straight through a fixed question. Probing absorbs some of it by pursuing each block's objective until it is met.
Three limits on this reading:
- Ceiling compression. Brand perception sits at 5.88 of 6, where the SD is squeezed mechanically. Purchase journey (7.29 of 9) is the cleanest row. The paper tests means (Welch) and reports no variance test, so these are descriptive figures, not a significance result.
- The coder scores against the AI's own brief. The breadth codebook is the moderator objective broken into parts, and only the AI was told to keep pursuing it (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins, measurement caveats). Lower AI dispersion on that metric partly reflects the AI being instructed to fill in that checklist. The length-controlled regressions and the blind pairwise judge point the same way, but none of them uses an independent codebook.
- Market research is not hiring. Respondents have no stake in the outcome, and nothing is decided about them.
How it maps onto the hiring result#
The hiring paper cannot run this split, because it has no static arm. It does contain one bridge. In the within-human-arm regression that links applicant language to offers, number of exchanges is one of the positive predictors, and the AI arm scores higher on it (Controlled Variance: AI's Edge as Reduced Dispersion, mechanism section; raw §4.3). Number of exchanges is the transcript feature that most directly records probing. In Deng et al. the adaptive arm ran 30.5 turns against the static arm's 8.3. That is consistent with the same ingredient operating in hiring. It is not evidence that it does, because the offer regression is correlational and confined to the human arm.
The result is a better name for the information-side channel, not a size for it. What is attributed to "information collection" in the +12% is best described as adaptive probing inside a guide that is followed consistently. "Standardization" is the wrong description, because the randomized evidence shows standardization on its own is the weakest configuration. There are two tensions with how Controlled Variance: AI's Edge as Reduced Dispersion states the concept, and both are refinements rather than contradictions:
- Its summary says "AI wins by being less dispersed, not more capable." The randomized evidence says that low procedural dispersion without adaptivity loses. What the AI arms share is low outcome dispersion that the adaptivity produces.
- Deng et al.'s twin result is that 56% more words and 3.7× more turns added no predictive signal (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins). Richer collected information is therefore not automatically more decision-relevant information. This fits the hiring paper's finding that recruiters discounted the AI interview score, which became less predictive of offers while the language test became more predictive (Controlled Variance: AI's Edge as Reduced Dispersion). The information-collection channel runs through what recruiters actually use, and neither study measures that directly.
2. Why the abort side does not move#
Deng et al. have no abort channel at all. Screening happens in a Section 0 screener on the platform, or at a Qualtrics intake for the human arm, before any interview starts. That is a design lesson, not evidence: a hiring replication that moved the non-negotiable-requirement screen (location, visa, rehire status) into an identical pre-interview screener in both arms would remove the abort channel by construction. Whatever treatment effect remained would then be pure information collection. The existing proposals, a completion-conditional re-estimate or an instrument for the screen-out decision, try to recover the split after the fact. This design prevents it from arising.
The other hiring-adjacent source in the vault, Procedural Value in AI Decisions, varies who decides in a stated-preference conjoint. It measures no in-interview screen-outs. The three things that kept the question open on 2026-08-17 are still unaddressed by any source: (i) r, the offer rate the differentially screened-out applicants would have had; (ii) per-arm evaluable-record rates over the randomized denominator; (iii) whether the hiring agent had an abort action at all.
3. The unidentified system#
Deng et al. identify their stack fully: GPT-5.1 at temperature 0 with 300-token turns, gpt-4o-transcribe in, tts-1 out, and a turn-based, goal-conditioned design that "generated either a clarification or follow-up within the current section, or a transition to the next section". Against expert human moderators, its breadth SDs are lower in all three reported blocks (0.91 vs 1.61, 0.80 vs 0.88, 0.35 vs 1.51). That is the controlled-variance signature from a different system, vendor and task. Two caveats keep this weak. The AI-vs-human contrast is not randomized (separate later sample, intake survey, live scheduling, $12 vs $10, n=24). And the hiring agent is still unnamed, so a replication still cannot know what it is replicating. The question moves slightly on generality and not at all on identification.
Verdict#
- Information vs abort (
#oq/now): partially answered, stays open. What moved: the information side now has randomized evidence, from market research, that adaptive probing produces both the content gain and the lower outcome dispersion, and that zero-discretion standardization produces neither. The channel should be called adaptive probing within a consistent guide, not standardization. What remains: any measurement of the abort channel's share. The 2026-08-17 bounds (+16% symmetric, r* = 5.8%) still stand as bounds, not an identified decomposition. A pre-interview screener common to both arms is the cleanest design that would settle it. - Unidentified system (
#oq/source): partially answered. A named second system shows the lower-dispersion signature against humans, which is weak and confounded evidence that controlled variance is not specific to one agent. The hiring study's system remains unidentified.
Sources#
- Controlled Variance: AI's Edge as Reduced Dispersion — the hiring RCT: procedural-consistency table, the eight-feature applicant-language regression (number of exchanges as a positive offer predictor), recruiter signal discounting, the 25%/7% screen-out split, and both open questions.
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins — the market-research RCT: design, the randomized AI-vs-static contrast and the non-randomized human arm, Table 6/7 results, measurement caveats (objective-derived codebook, unvalidated coder), the twin null, and the scope note that it has no screen-out function.
- What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators — the abort-channel bounds this page leaves standing.
- Procedural Value in AI Decisions — checked as the only other hiring-procedure source; it varies decision authority, not in-interview screening.
- AI-Moderated Interviews for Market Research and Digital Twins Calibration — §2.1 (design: randomization of AI and static arms in one batch; static arm "fixed, pre-defined questions with no adaptive follow-up or probing"; GPT-5.1 stack; human arm recruited separately), Table 7 (breadth means and SDs by block and arm, the source of the CV column), Appendix C (objective-derived codebook), Appendix G (length controls).
- Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — §4.3 two-step linguistic analysis (number of exchanges defined as interviewer-then-applicant sequences, Appendix variable definitions), §4.1 interview-type classification.
Cited by 5
- Controlled Variance: AI's Edge as Reduced Dispersion×3
The AI system is never identified — no model, vendor, or version, only that Google Cloud supplied…
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins
Adaptive Probing Vs Standardization In Ai Interviews — uses this page's randomized AI-vs-static…
- AI Economics & Labor
Adaptive Probing Vs Standardization In Ai Interviews — Re-works the open question on Jabarian &…
- Open Questions Dashboard
Controlled Variance: How much of the +12% is controlled variance in information collection versus…
- Open Questions Backlog
Controlled Variance: The AI system is never identified — no model, vendor, or version, only that…
Related articles
- Controlled Variance: AI's Edge as Reduced Dispersion
Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants betwe…
- Configurable Human Participation
HAS-Bench (Wu et al.): human participation as a configurable benchmark variable (five-level agency scale × three intera…
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins
Deng, Liu, Toubia & Jain (arXiv 2609.29143): a preregistered market-research study (GPT-5.1 voice moderator N=139 vs sa…
- The Automation–Optimism Link
AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-qualit…
- Diversity Calibration Under SFT
Skobelev, Fithian & Han (arXiv 2609.16454): output diversity is measured as collision probability, and plain SFT is not…
