H
Howardism
Plate IIAI Economics & LaborHOWARDISM

Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect

Re-works the open question on Jabarian & Henkel's +12% AI-interviewer offer effect (information collection vs the interviewer's discretion to abort) with the vault's second AI-interviewer study, Deng, Liu, Toubia & Jain's market-research RCT. Its randomized AI-vs-static arms split the information side: the zero-discretion static guide is the weakest arm on every content measure, and on this page's arithmetic from Table 7 it is also the more dispersed in outcome (breadth CV 0.24 vs 0.12, 0.31 vs 0.24, 0.13 vs 0.06), so a uniform procedure does not give uniform information. The working ingredient is adaptive probing inside a fixed guide. The abort side does not move. No vault source measures the screen-out channel, and the 2026-08-17 bounds (+16% symmetric, r* = 5.8%) are still the best available. Also a named second system (GPT-5.1) showing the lower-dispersion signature against humans, confounded and n=24, which is weak evidence that the property is not specific to one agent

Article metadata
Publication details
Published:October 1, 2026
Filed:Essay
Domain:AI Economics & Labor
Reading:11 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect

Sources#

The question#

From Controlled Variance: AI's Edge as Reduced Dispersion (#oq/now, partially answered 2026-08-17 by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators): how much of the +12% is controlled variance in information collection versus removal of the interviewer's discretion to abort? This page also covers the sibling #oq/source question on the same page: the AI system is never identified, so is "controlled variance" a property of this one agent or of AI-conducted interviews generally?

The new input is AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins (Deng, Liu, Toubia & Jain, arXiv 2609.29143, empirical, preregistered). It randomizes an adaptive GPT-5.1 voice moderator (N=139) against a static arm with the same fixed questions (N=154), on the same platform and in the same recruitment batch. A human-moderator arm (N=24) was recruited separately and later.

Short answer#

The question has three parts, not two. Jabarian & Henkel's AI arm differs from human recruiters in three ways at once. It follows the guide more uniformly (standardization). It still tailors follow-ups to each applicant (adaptivity). And it screens out far less often (abort: 7% vs 25%). The 2026-08-17 synthesis bounded the abort part. The new study is randomized evidence on the standardization-vs-adaptivity split, from a different task.

  1. Information side: adaptivity carries it and standardization alone does not. Deng et al.'s static arm is pure standardization: identical questions, no probing, no interviewer discretion, no aborts. It is the weakest arm on every content measure: 5.01 vs 7.48 needs per interview, and lower breadth in three of four blocks, which survives length control (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins). It is also the more dispersed arm in outcome, as shown below. So the paper's own description of the mechanism, "consistent protocol, tailored follow-ups", is doing the work through its second half.
  2. Abort side: unchanged. Deng et al. settle eligibility with screeners before the interview, neither moderator could end an interview on a disqualifier, and the paper records no aborts. No other vault source measures a screen-out channel. The bounds on What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators remain the best available: the arms lose about the same share of interviews (43% vs 45%), the symmetric completion-conditional correction gives +16%, and the break-even is r* = 5.8%. The decomposition itself still has no direct measurement.
  3. Sibling question: weak movement on generality, none on identification. A second, named system (GPT-5.1, temperature 0, on an unnamed startup's platform) shows lower coverage dispersion than expert human moderators in a second task. That is a small amount of evidence that the signature is not idiosyncratic to Jabarian & Henkel's agent. It is on the non-randomized contrast with n=24, and it does nothing to identify the hiring study's system.

1. What the randomized split shows#

The static arm is least discretionary and most dispersed#

Controlled Variance: AI's Edge as Reduced Dispersion measures the hiring AI's advantage in procedural terms: topic-order τ 0.53 vs 0.33, question-to-guideline similarity 0.59 vs 0.43, and lower cross-interview variance on both. A reading of "controlled variance" that takes it literally would predict that the most uniform procedure yields the most uniform information. Deng et al.'s static arm is the limiting case. Every participant gets exactly the same questions, so procedural dispersion is zero by construction.

Breadth dispersion by arm, from Table 7 of the raw (mean breadth with SD in parentheses; the table was verified cell by cell against pdftotext at compile time, see AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins Sources). The coefficient of variation (CV) is this page's arithmetic.

Block (max breadth)AI-moderatedStaticAI CVStatic CV
Purchase journey (9)7.29 (0.91)5.90 (1.44)0.120.24
Brand awareness (4)3.32 (0.80)1.73 (0.53)0.240.31
Brand perception (6)5.88 (0.35)5.52 (0.69)0.060.13
Category usage (3, AI capped at 3 questions)2.94 (0.25)2.94 (0.24)0.090.08

In the three blocks where the AI could probe freely, the adaptive arm has the lower relative dispersion in all three and the lower raw SD in two. In the block where probing was capped, the arms are identical. So removing all interviewer discretion did not produce uniform information. The respondents' own variation passes straight through a fixed question. Probing absorbs some of it by pursuing each block's objective until it is met.

Three limits on this reading:

  • Ceiling compression. Brand perception sits at 5.88 of 6, where the SD is squeezed mechanically. Purchase journey (7.29 of 9) is the cleanest row. The paper tests means (Welch) and reports no variance test, so these are descriptive figures, not a significance result.
  • The coder scores against the AI's own brief. The breadth codebook is the moderator objective broken into parts, and only the AI was told to keep pursuing it (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins, measurement caveats). Lower AI dispersion on that metric partly reflects the AI being instructed to fill in that checklist. The length-controlled regressions and the blind pairwise judge point the same way, but none of them uses an independent codebook.
  • Market research is not hiring. Respondents have no stake in the outcome, and nothing is decided about them.

How it maps onto the hiring result#

The hiring paper cannot run this split, because it has no static arm. It does contain one bridge. In the within-human-arm regression that links applicant language to offers, number of exchanges is one of the positive predictors, and the AI arm scores higher on it (Controlled Variance: AI's Edge as Reduced Dispersion, mechanism section; raw §4.3). Number of exchanges is the transcript feature that most directly records probing. In Deng et al. the adaptive arm ran 30.5 turns against the static arm's 8.3. That is consistent with the same ingredient operating in hiring. It is not evidence that it does, because the offer regression is correlational and confined to the human arm.

The result is a better name for the information-side channel, not a size for it. What is attributed to "information collection" in the +12% is best described as adaptive probing inside a guide that is followed consistently. "Standardization" is the wrong description, because the randomized evidence shows standardization on its own is the weakest configuration. There are two tensions with how Controlled Variance: AI's Edge as Reduced Dispersion states the concept, and both are refinements rather than contradictions:

  • Its summary says "AI wins by being less dispersed, not more capable." The randomized evidence says that low procedural dispersion without adaptivity loses. What the AI arms share is low outcome dispersion that the adaptivity produces.
  • Deng et al.'s twin result is that 56% more words and 3.7× more turns added no predictive signal (AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins). Richer collected information is therefore not automatically more decision-relevant information. This fits the hiring paper's finding that recruiters discounted the AI interview score, which became less predictive of offers while the language test became more predictive (Controlled Variance: AI's Edge as Reduced Dispersion). The information-collection channel runs through what recruiters actually use, and neither study measures that directly.

2. Why the abort side does not move#

Deng et al. have no abort channel at all. Screening happens in a Section 0 screener on the platform, or at a Qualtrics intake for the human arm, before any interview starts. That is a design lesson, not evidence: a hiring replication that moved the non-negotiable-requirement screen (location, visa, rehire status) into an identical pre-interview screener in both arms would remove the abort channel by construction. Whatever treatment effect remained would then be pure information collection. The existing proposals, a completion-conditional re-estimate or an instrument for the screen-out decision, try to recover the split after the fact. This design prevents it from arising.

The other hiring-adjacent source in the vault, Procedural Value in AI Decisions, varies who decides in a stated-preference conjoint. It measures no in-interview screen-outs. The three things that kept the question open on 2026-08-17 are still unaddressed by any source: (i) r, the offer rate the differentially screened-out applicants would have had; (ii) per-arm evaluable-record rates over the randomized denominator; (iii) whether the hiring agent had an abort action at all.

3. The unidentified system#

Deng et al. identify their stack fully: GPT-5.1 at temperature 0 with 300-token turns, gpt-4o-transcribe in, tts-1 out, and a turn-based, goal-conditioned design that "generated either a clarification or follow-up within the current section, or a transition to the next section". Against expert human moderators, its breadth SDs are lower in all three reported blocks (0.91 vs 1.61, 0.80 vs 0.88, 0.35 vs 1.51). That is the controlled-variance signature from a different system, vendor and task. Two caveats keep this weak. The AI-vs-human contrast is not randomized (separate later sample, intake survey, live scheduling, $12 vs $10, n=24). And the hiring agent is still unnamed, so a replication still cannot know what it is replicating. The question moves slightly on generality and not at all on identification.

Verdict#

  • Information vs abort (#oq/now): partially answered, stays open. What moved: the information side now has randomized evidence, from market research, that adaptive probing produces both the content gain and the lower outcome dispersion, and that zero-discretion standardization produces neither. The channel should be called adaptive probing within a consistent guide, not standardization. What remains: any measurement of the abort channel's share. The 2026-08-17 bounds (+16% symmetric, r* = 5.8%) still stand as bounds, not an identified decomposition. A pre-interview screener common to both arms is the cleanest design that would settle it.
  • Unidentified system (#oq/source): partially answered. A named second system shows the lower-dispersion signature against humans, which is weak and confounded evidence that controlled variance is not specific to one agent. The hiring study's system remains unidentified.

Sources#

§ end
Cited by 5
Related articles