H
Howardism
Plate IIAI Economics & LaborHOWARDISM

What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators

PublishedAugust 17, 2026FiledEssayDomainAI Economics & LaborReading28 minSourceAI-synthesised

Answers two open questions as one: the +12% AI-interview offer effect and ATLAS's 22.6% classifier accuracy are both ratios whose denominator was assumed rather than measured. For the first, the paper's own figures bound the abort channel — the arms lose almost the same share of interviews (43% human vs 45% AI), so the *symmetric* completion-conditional correction moves the effect **up** to +16%, not down; only the one-sided correction the question proposes flips the sign (−10%), and the screen-out variable it would run on is an unvalidated LLM label whose codebook thresholds on a variable the treatment moves. For the second, yes: leave-one-annotator-out human agreement is the ceiling estimator (ATLAS has the annotations and never computes it), and renormalizing gives 73–82% of the human ceiling at occupation title and 85–94% at major group — while at O*NET-task granularity ATLAS proves no ceiling is estimable, which is why aggregating categories *is* the ceiling fix

Illustration for What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators

Sources#

The questions#

  1. From Controlled Variance: AI's Edge as Reduced Dispersion: How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them.

  2. From Usage-Telemetry Classifier Validation: Human raters disagree with each other on 42–48% of 3-digit occupations. Is there a principled way to establish the ceiling a classifier could reach, so accuracy can be reported relative to it rather than to 100%?

They are one problem. Both headline numbers are ratios reported against a denominator the study asserted rather than measured: +12% against all randomized applicants, which silently pools two stages (reaching the evaluation, and winning at it); 22.6% against 100%, which assumes a rater exists who could reach it. Q1 asks what happens when you fix the denominator; Q2 asks how you would measure one.

Short answer#

Q1 — the abort channel almost certainly is not the source of the effect, and the correction the question proposes is the one that would make it look like it is. The two arms lose nearly the same share of interviews to early termination (43% human vs 45% AI); what differs is the composition — the human arm's losses are recruiter-initiated screen-outs selected on fit, the AI arm's are applicant- or machine-initiated failures that are not. Correcting only the human screen-outs flips the effect to −10%; correcting only the AI's own 12% technical-failure-plus-refusal channel lifts it to +27%; correcting both symmetrically gives +16%, i.e. the effect survives and grows. The paper's figures bound the channel but cannot decompose it, because the two quantities required — per-arm evaluable-record rates over the randomized denominator, and the offer rate among early-terminated interviews — are never published, and the screen-out variable itself is an unvalidated LLM label.

Q2 — yes, and the wiki documents the estimator: leave-one-annotator-out agreement, computed on the same items with the same protocol. ATLAS already collected the annotations that would produce it (three raters, N ≈ 110–120) and instead borrowed a ceiling from 1983/1992 survey literature. Renormalizing its accuracy against that borrowed ceiling turns 42.5% at occupation title into 73–82% of the human ceiling and 71.6% at major group into 85–94%. At O*NET-task granularity no ceiling is estimable at all — ATLAS proves this itself — which reframes its category-merging mitigation: merging categories humans also cannot separate is not noise reduction, it is raising the ceiling.


Part 1 — The abort channel, bounded by the paper's own arithmetic#

What the paper reports#

All figures from the paper's prose (§3.1 and §4.1) — see the parse note below.

QuantityHumanAI
Randomized applicants13,55740,103
Offers1,179 = 8.70%3,904 = 9.73%
Screen-outs (share of interviews)25%7%
Comprehensive interviews39%42%
Disengaged9%12%
Applicant unavailable9%14%
AI technical abort7%
Applicant refused the AI5%
Sum of early-terminating types43%45%
Residual (expectation mismatch, other, non-English)18%13%

The last two rows are the finding the open question does not anticipate, and they are simple addition on figures Controlled Variance: AI's Edge as Reduced Dispersion already carries. The arms lose nearly the same fraction of interviews before a normal completion. The AI's 18pp deficit in screen-outs is almost exactly offset by its surpluses in disengagement (+3), unavailability (+5), technical failure (+7) and refusal (+5). So the intervention did not reduce the volume of aborted interviews. It changed who decides to abort and on what information: the human arm's terminations are recruiter-initiated and conditioned on a stated disqualifier, the AI arm's are applicant-initiated or mechanical and conditioned on nothing about fit.

That is a materially different claim from either "AI collects better information" or "we removed the interviewer's option to give up early," and it is the one the reported numbers actually support.

The three one-sided corrections#

Treat early-terminated interviews as producing no offer, and re-express the offer rate over the surviving share. Three corrections are available; they disagree violently, and which one you run is the whole question.

CorrectionHumanAIEffect
None (published ITT)8.70%9.73%+1.03pp, +12%
Remove human screen-outs only (the question's proposal)8.70/0.75 = 11.60%9.73/0.93 = 10.47%−1.13pp, −10%
Remove AI technical failures + refusals only8.70%9.73/0.88 = 11.06%+2.37pp, +27%
Remove all early-terminating types, both arms8.70/0.57 = 15.26%9.73/0.55 = 17.70%+2.44pp, +16%

Three points follow.

The question's proposed re-estimate is asymmetric, and that asymmetry is what produces the sign flip. Conditioning on "the human recruiter did not screen this applicant out" while leaving the AI arm's 12% of interviews that crashed or were refused inside the denominator corrects one arm's aborts and not the other's. Applied to the AI arm's own failure channel, the identical move more than doubles the effect. Neither number is an estimate; they are the two ends of a one-sided sensitivity analysis.

The symmetric correction is the defensible one, and it moves the effect up. Conditioning both arms on reaching a normal completion — the closest available implementation of "the subsample of interviews that reached completion in both arms" — gives +16%, larger than the published +12%. The mechanism is the composition point above: the human arm's aborts remove applicants selected for failing a requirement, whose counterfactual offer probability is low, so removing them from the denominator inflates the human rate a lot while adding little that would have been an offer; the AI arm's aborts remove applicants selected for being unreachable, uninterested, or unlucky with the software, whose counterfactual offer probability is closer to average, so the AI arm was paying a genuine price for them in the ITT. The paper says as much informally — the +12% is net of a 12% in-arm failure rate (Controlled Variance: AI's Edge as Reduced Dispersion) — but never carries it through to the arithmetic.

Every one of these conditions on a post-treatment variable and is therefore not a causal estimate of anything. They bound how much room the channel has; they do not identify it. The question's second proposal — instrument the screen-out decision — remains the only design that would.

The break-even bound#

The cleanest single number the published figures permit. Let r be the offer rate the differentially screened-out applicants would have achieved had their interviews run to completion. The screen-out share differs by 18pp, so the abort channel contributes 0.18 × r percentage points to the human arm's deficit. Setting that equal to the measured 1.03pp:

r* = 1.03 / 18 = 5.8% — which is 66% of the human arm's own base offer rate of 8.70%.

So the abort channel accounts for the entire treatment effect if and only if the applicants human recruiters screened out, but the AI would not have, were about two-thirds as likely to receive an offer as the average applicant. The share of the effect attributable to the channel is r / 5.8%: at r = 1% it is 17%, at r = 3% it is 52%, at r = 8.7% it is 152%.

Is r plausibly that high? The paper's own codebook argues no. The three screen-out types are defined by disqualification on a non-negotiable requirement — the worked examples are location, salary expectation, visa status, conflicting school plans, and rehire ineligibility (Appendix E.1 and the classifier prompt in J.3). Applicants failing a non-negotiable firm requirement have a counterfactual offer probability near zero by construction, which puts r far below 5.8% and the channel's share far below 100%. The counter-argument that keeps the question open: Late Screen-Out is defined as an interview that "proceeds nearly to completion" with topic count ≥ 8 before failing a final criterion, which describes an applicant who looked hirable for most of the conversation. The paper reports only the pooled 25% / 7%, never the early/midway/late split, so the share of screen-outs that are late — the high-r ones — is unpublished.

The two unpublished quantities#

The decomposition needs exactly two things the paper does not report, and no amount of arithmetic on what it does report substitutes for either:

  1. Per-arm evaluable-record rates over the randomized denominator. The 25% / 7% shares are computed over interview transcripts — 34,109 for both arms combined, explicitly "a subset of all interviews conducted," against 53,660 randomized applicants in those arms. Every correction above assumes the transcript sample is representative and that the two arms convert randomized applicants into conducted interviews at the same rate. The second assumption is doubtful in a known direction: the AI agent is available 24/7 and cut median time-from-interest-to-interview from 0.51 to 0.32 days (Controlled Variance: AI's Edge as Reduced Dispersion), so the AI arm plausibly interviews more of the people assigned to it. That difference sits entirely inside the ITT denominator and is invisible in the published tables.
  2. The offer rate among early-terminated interviews, by arm and termination type. The firm has it — offers are administrative — and it is what pins r.

The deeper problem: the variable the decomposition would run on is an unvalidated LLM label#

This is where Q1 collides with Q2, and it is the reason the answer is partial rather than merely unfinished.

The 25% / 7% screen-out split is not an administrative field. It is the output of an LLM classifier: raw transcripts preprocessed by gemini-2.0-flash for speaker tagging and PII removal, then sorted into ten mutually exclusive interview types by a prompt combining role-based framing, chain-of-thought and in-context examples (Appendix E.1, prompt in J.3). The paper reports no accuracy figure, no human agreement statistic, no κ, and no error analysis for it. It is precisely the instrument Usage-Telemetry Classifier Validation exists to interrogate, load-bearing for the mechanism half of the paper, and validated to the reader in general terms only — the same public asymmetry that page already flags for the AEI's and OpenAI's classifiers.

Two specific defects make the labels worse than merely unvalidated:

The codebook requires an utterance the AI may structurally never produce. All three screen-out types carry the rule "Recruiter states the reason for ending the call due to disqualification," reinforced by "if a call ends without concluding remark from the recruiter, [Screen-Out] does not apply." The paper never says whether the AI voice agent was given the authority or the instruction to terminate an interview on a disqualifier — only that it was "prompted to follow the same structured interview guidelines." If it was not, then 7% is not a measurement of an agent declining to abort; it is a measurement of an agent that had no abort action, and the events that would have been screen-outs land in Disengaged, Unavailable, or Other. The residual category is 18% in the human arm against 13% in the AI arm, so the missing mass is not parked there, but the label's dependence on a spoken disqualification is a treatment-coupled definition either way.

The codebook thresholds on a variable the treatment moves. Disengaged Interaction is defined as topic_count < 8; Comprehensive Interview as topic_count ≥ 8; the three screen-out tiers are separated by topic-count bands (0–2 / 3–7 / ≥8). The treatment raises mean topic coverage from 38% to 45% (Controlled Variance: AI's Edge as Reduced Dispersion). A rightward shift in the coverage distribution therefore moves probability mass across these thresholds mechanically. The +3pp shift in comprehensive interviews (39% → 42%, and +2.99pp / +4.82pp in the paper's own Table B.10) is largely a discretization of the +7pp coverage shift, not independent corroboration of it — and the early/midway/late screen-out mix is moved by the same arithmetic.

This is LLM-Judge Validation's missing-step problem in an economics paper rather than an eval harness: the classifier is stratified by nothing, validated against nothing, and its decision rule is correlated with the treatment. Answering Q1 properly requires validating it first, which makes Q2 a precondition for Q1 rather than a parallel question.


Part 2 — Yes, there is a principled ceiling, and ATLAS has the data to compute one#

The estimator: leave-one-annotator-out agreement on the same items#

The principled construction is not new, and the corpus already contains a worked instance. In LLM-Judge Validation, Yang et al. calibrate judges against a leave-one-annotator-out human ceiling: hold out one annotator, predict their label from the others, and score the model by the identical procedure on the identical items. The verdict is a ratio rather than a raw number — PandaLM's best judge reaches κ = 0.753 against a human ceiling of 0.920 (headroom remains), while Judge's Verdict's best judge reaches κ = 0.620 against a human ceiling of 0.562, i.e. the judges beat the humans, which is a statement about the labels rather than the judges.

That is exactly what Q2 asks for, and ATLAS already collected the ingredients. Its Appendix B.3.1 runs three in-house annotators over N ≈ 110–120 clusters and reports, for SOC Major Group, the model's Cohen's κ against human plurality consensus (0.83, Manski bounds [0.75–0.85]) alongside human-only pairwise κ (0.66) and human-only Fleiss κ (0.68). The comparison is one step short of a ceiling:

  • Model-vs-plurality-of-three is an easier target than human-vs-human pairwise. Agreeing with the consensus of a panel is not the same task as agreeing with one randomly drawn rater, and reading "0.83 > 0.68, the model beats the humans" off those two columns is the comparison the leave-one-out design exists to prevent. ATLAS is careful in its prose ("performs comparably to (or may be even better) than a human"); the number invites the stronger reading.
  • The fix costs no additional annotation. With three raters per cluster the held-out design is computable from the data already collected: score each human against the plurality of the other two, then score the model against the same two-rater plurality. That single recomputation converts ATLAS's incomparable pair into the ratio Q2 wants.

The second instrument, from the same page: LLM-Judge Validation's Shopify Sidekick account estimates the ceiling before the classifier exists, on the actual rubric, by the actual annotators, on the actual traffic — two experts blind-annotate 25 production samples, agreement is recorded as Cohen's κ, and κ ≈ 0.2 is a rubric-rewrite trigger rather than a pass mark. The judge is then aimed at that number instead of at 100% ("the goal is not a judge that is 'perfect,' but one that matches humans about as well as humans match each other"; reported human agreement 83% against judge 80%, with the chart annotating perfection as "unreachable"). The contrast with ATLAS is the point: ATLAS's ceiling is borrowed from the literature (Mellow & Sider 1983, Mathiowetz 1992), Shopify's is measured on the task. A borrowed ceiling tells you a ceiling exists; a measured one tells you where to stop optimizing.

The general form of the move is on Matched Comparisons for Memorization Claims: correct a raw rate by a baseline the field had been assuming was zero. There the assumed-zero quantity is the generation rate on non-training data; here it is the human error rate. Both corrections are large, and in both the raw number is uninterpretable without them.

The renormalization, worked#

Using ATLAS's own borrowed ceilings — worker-employer disagreement of 17% / 24% at 1-digit and 42% / 48% at 3-digit occupation — and reporting (accuracy − chance) / (ceiling − chance), the chance-corrected share of the achievable range:

LevelCategoriesChanceAccuracyHuman ceilingRenormalized
SOC Major Group234.35%71.57%76–83%85–94%
SOC Occupation Title1,0160.10%42.47%52–58%73–82%
O*NET Specific Tasks18,7970.005%22.58%none existsnot computable

So the sentence "the classifier assigns the exact occupation title correctly 42.5% of the time" becomes "the classifier recovers roughly three-quarters to four-fifths of the agreement two humans reach on occupation coding at comparable granularity" — a different claim about the same pipeline, and the one that bears on whether the economics is usable.

Three caveats, all of which say the same thing: this is an illustrative renormalization, not a validated one, and it exists to show what the reported form would look like rather than to supply the number.

  • Granularity mismatch, in a known direction. 1,016 SOC occupation titles is finer than the 3-digit census occupation the 42–48% disagreement was measured on, and 23 SOC major groups is coarser than the 11-category 1-digit taxonomy of Mellow & Sider's era — a mismatch ATLAS flags itself. Both borrowed ceilings are therefore too high for the level they are applied to, so both renormalized figures understate the classifier's share of the achievable range.
  • Population and task mismatch. The accuracy figures come from 18,801 synthetic, Gemini-generated conversations; the ceilings come from 1980s–90s US household surveys where a worker and an employer independently described a job. Different raters, different task, different evidence, different taxonomy vintage. Dividing one by the other is a sanity check, not an estimate — which is precisely the argument for the leave-one-out design, where ceiling and score come from the same items.
  • The synthetic-data caveat cuts both ways and ATLAS states it: a Gemini-generated prompt classified by a Gemini classifier may carry recoverable cues (inflating accuracy), or an ambiguous prompt may genuinely fit a category better than its seeded label (deflating it).

The approval number is not a ceiling, and cannot be made into one#

The tempting shortcut — treat the 85.8% human approval rate as the ceiling and read 22.6/85.8 — is wrong, for a mechanism the corpus measures precisely.

Approval asks a rater "is this label defensible?" with the label shown. Blind coding asks them to produce one. ATLAS's own data prices the difference: rater approval exceeds model-versus-plurality agreement by 11–15pp at the top taxonomy levels, and the report names the cause ("raters [can] pick different options unprompted but still agree with an alternative if one is presented"). Reference-Free Judge Over-Crediting measures the same effect at its extreme in the judge setting — requiring the rater to commit its own answer before seeing the candidate collapses the false-positive rate from 0.719 to 0.012 on identical text — and LLM-Judge Validation already names the human-side version as a missing control: "have a subset of annotators label blind, then reveal the judge's verdict, and report both numbers."

So the two numbers Q2 juxtaposes are on different scales by construction: 42–48% disagreement is a blind statistic (two parties independently describing the same job) while 85.8% approval is an anchored one. Approval is an upper bound of unknown tightness on validator-classifier agreement, not a ceiling on classification.

Where the ceiling is provably unmeasurable — and why that is the answer, not a gap in it#

ATLAS declines to report inter-rater agreement below ATUS Tier 2 / SOC Minor, and gives a statistical reason (B.3.1) that generalizes to any attempt at the O*NET task level: when the number of candidate categories vastly exceeds the number of rated observations, the hundreds or thousands of unsampled categories contribute nothing to expected chance agreement, so p_e is overestimated and κ underestimated; worse, if agreement varies by category, a sample covering a tiny fraction of categories carries no information about the population, and the computed κ is driven by whichever categories happened to be drawn.

That is a proof, not a budget constraint. No feasible annotation study establishes a human ceiling over 18,797 O*NET task statements, so the 22.58% has no reportable denominator and will not acquire one by spending more on raters. Two consequences:

  1. The honest report at that granularity is what ATLAS actually says — that approval and accuracy are bounds on confidence, and the exact-assignment figure is "a caution against over-relying on hyper-specific task analysis." A renormalized number would be fabricated.
  2. Aggregation is the ceiling intervention, not a noise-reduction trick. ATLAS's three published mitigations look like error management and read differently through this frame. Mapping exact tasks into Autor–Thompson task types takes 22.58% → 70.44% at a level where 5 categories and 20% chance make an agreement study feasible. Merging categories the taxonomy splits for non-economic reasons — travel-by-purpose, household versus non-household caregiving — lifts ATUS Tier 2 from 42.9% to 51.7% and Tier 3 from 23.7% to 32.4%. Those merges target exactly the distinctions users have no economic reason to disclose, which is to say the distinctions a human rater reading the same conversation could not make either. Merging them raises the ceiling and the accuracy together. The ceiling is a property of the category system, not of the classifier — so "report accuracy relative to a ceiling" and "merge categories nobody can separate from the available evidence" are the same intervention seen from two sides, and only the first tells you when to stop.

The reporting standard this implies#

For any taxonomy classifier whose output carries an economic claim:

  1. Report chance-corrected agreement (κ or α) as the headline, not exact match — LLM-Judge Validation's Finding 1, unchanged.
  2. Report a human ceiling measured on the same items by the same protocol, via leave-one-annotator-out. With three raters this needs no new annotation.
  3. Report accuracy as a share of the achievable range, (observed − chance) / (ceiling − chance).
  4. Report blind and anchored validator numbers separately; never use an approval rate as a ceiling.
  5. Where the ceiling is not estimable at the reporting granularity, say so and aggregate to the level where it is — and treat the aggregation as raising the ceiling rather than as hiding error.

Steps 1, 4 and 5 ATLAS effectively does. Step 2 it has the data for and does not run. Step 3 nobody in the corpus runs.


Part 3 — The one problem, stated once#

Both questions are instances of a single failure: a headline ratio whose denominator was assumed instead of measured, where the assumption is invisible because the denominator looks like a natural constant.

  • "All randomized applicants" looks like the neutral denominator for a randomized experiment, and is — for the ITT. It stops being neutral the moment the claim shifts from what happens if you assign an AI interviewer to why, because the offer rate then pools two stages the design deliberately separated, and the arms do not reach the second stage in the same way.
  • "100%" looks like the neutral denominator for an accuracy figure. It stops being neutral as soon as the labels are ones humans disagree about 42–48% of the time.

The corrective is the same in both: measure the denominator on the same sample with the same instrument. For Q1 that is the completion-conditional subsample — which the firm has and the paper does not report. For Q2 it is the leave-one-out human ceiling — which ATLAS collected and does not compute. And in both cases the correction is itself gated on validating an LLM classifier, which makes Q2's answer the precondition for Q1's.

This is also the sharper form of a rule Telemetry vs. Survey Measurement already carries — weight instruments by what they can see. Randomization sits a rung above telemetry and survey because it identifies a counterfactual, and the Jabarian & Henkel design is the vault's only instance of it. But the rung buys identification of the assigned contrast only. Everything past that — which channel, which stage, why — runs on the same observational, classifier-mediated evidence as the telemetry studies, with the same unquantified error bar. A randomized experiment does not upgrade the instruments inside it.


Verdict on the two questions#

Q1 — partially answered; stays open. What is settled: the abort channel is bounded, the symmetric completion-conditional correction moves the effect up to +16% rather than down, and the sign flip to −10% is an artifact of correcting one arm's aborts and not the other's. The r* = 5.8% break-even, against a codebook defining screen-outs by non-negotiable disqualification, makes the channel unlikely to carry the effect. What is not settled and cannot be from published data: r itself; the early/midway/late screen-out split; per-arm evaluable-record rates over the randomized denominator; and whether the AI agent possessed an abort action at all. The instrument the whole decomposition would run on — a ten-way LLM interview-type classifier — has no published validation and thresholds on a variable the treatment moves.

Q2 — answered. Yes: leave-one-annotator-out agreement on the same items, with accuracy reported as (observed − chance) / (ceiling − chance); ATLAS's own annotations already support it. Worked renormalization gives 73–82% of the human ceiling at occupation title and 85–94% at major group, with the borrowed-ceiling caveats above. Where the category count defeats any feasible agreement study — O*NET's 18,797 tasks — no ceiling is estimable, which is itself part of the answer: the available move is to aggregate to a level where it is, and that aggregation is the ceiling fix.


Parse notes#

Per the PDF-table discipline in _system/compiler-prompt.md, every number above is quoted from prose unless noted, and each table row cited was reconciled first.

  • Jabarian & Henkel. All offer, retention, screen-out and interview-type figures come from §3.1 and §4.1 prose; the codebook definitions from Appendix E.1 and the verbatim classifier prompt in J.3. Table B.10 (treatment effect on "interview is comprehensive") was used only as corroboration and carries a table-collapse: two label rows are welded into one (Mean DV in Human Interviewer Controls and fixed effects | 0.3879 - | 0.3880 Yes). The values are recoverable and reconcile with prose — human mean 0.3879 plus the +0.0299 coefficient gives 41.8%, matching the prose's 39% / 42% — so the row is quoted only in that reconciled form. The damaged Tables 1, B.5 Panel B and B.6 already documented on Controlled Variance: AI's Edge as Reduced Dispersion are not used here.
  • A definitional wrinkle worth knowing before converting the coverage percentages to topic counts: §4.2 describes 14 guideline topics, while Appendix E.2 defines topic coverage as covered topics divided by a maximum of 15 (the 14 topics plus an "other" bucket). The 38% / 45% figures are unaffected as percentages; the implied topic counts differ by ~7% depending on which denominator applies, which matters because the interview-type codebook thresholds on a raw count of 8.
  • ATLAS. Table 5 (accuracy by taxonomy level) parses clean — one row per level, distinct cells — and reconciles with prose ("nearly 72% accuracy at the SOC major group… stays above 57% even at the narrower SOC minor group"). Tables 6 and 7 parse clean, with the N column intact. One reconciliation matters for the argument: B.3.1's prose reports "a plurality agreement kappa of 0.83, and Fleiss' kappa of 0.69" for SOC Major, and 0.69 is the with-LLM Fleiss from Table 6, not the human-only figure, which is 0.68. Quoting the prose number as the human ceiling would place the model inside a ceiling that already includes it — the exact circularity the leave-one-out design removes.

Sources#

  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews, arXiv 2607.28222 (2026-07-30), empirical, pre-registered (AEA RCT #15385). §2.4–2.5 (design, arms, randomized counts, the AI agent's prompt described only as "the same structured interview guidelines"), §3.1 (offer / start / retention rates and counts, the conditional-on-acceptance sample), §4.1 (interview-type classification and the 25% / 7% screen-out split; transcript coverage of 34,109 applications as "a subset of all interviews conducted"), §4.2 (topic coverage 38% / 45% and the per-recruiter benchmark), §7 (time-to-hire decomposition), Appendix E.1 + Table E.1 (the ten interview types and their topic-count and duration bands), Appendix E.2 (topic coverage defined over 15 topics), Appendix J.1–J.3 (gemini-2.0-flash preprocessing; the interview-type classification prompt and its "recruiter states the reason" and "no concluding remark" rules). Full evidence note, disclosure and parse warnings at Controlled Variance: AI's Edge as Reduced Dispersion.
  • Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the EconomyATLAS v1.0, Appendix B: Classifier Validation. B.1 (Mellow & Sider 1983 at 17% / 42%, Mathiowetz 1992 at 24% / 48%; the ambiguity-predates-AI argument), B.2 + Table 5 (accuracy and chance rates by level; the 93.7% work/non-work gate; the synthetic-data caveat in both directions), B.2.2 (structured confusion; the +8.8pp / +8.7pp category-merge gains), B.3.1 + Table 6 (three annotators, N ≈ 110–120; model-vs-plurality κ, human-only pairwise and Fleiss κ; the explicit refusal to compute κ at finer levels and its reasoning), B.3.2 + Table 7 (the approval study, its single yes/no question, and the 11–15pp approval-over-plurality gap). Full treatment at Usage-Telemetry Classifier Validation and Google AI & Economy ATLAS.
  • Controlled Variance: AI's Edge as Reduced Dispersion — the experiment's full treatment: the +12% / +18% ITT results, the 12% in-arm AI failure rate, the bimodality qualification on the variance claim, the recruiter signal-discounting result, and the existing parse warnings.
  • Usage-Telemetry Classifier Validation — the accuracy/approval gulf, the aggregation and presence-not-frequency mitigations, and the statement of the ceiling problem this page answers.
  • LLM-Judge Validation — kappa deflation (exact match overstates κ by 33–41pp); Yang et al.'s leave-one-annotator-out human-ceiling calibration (PandaLM 0.753 against 0.920; Judge's Verdict 0.620 against 0.562); the Shopify Sidekick pre-judge ceiling protocol (blind expert pair, κ ≈ 0.2 rewrite trigger, 83% human against 80% judge); and the ratify-versus-predict asymmetry with its blind-then-reveal fix.
  • Reference-Free Judge Over-Crediting — the anchoring magnitude: commit-first de-anchoring collapses a judge's false-positive rate from 0.719 to 0.012 on identical text, which is why an approval rate cannot serve as a ceiling.
  • Matched Comparisons for Memorization Claims — the general form of the correction: a raw rate is uninterpretable until the baseline the field assumed was zero is measured by the identical procedure.
  • Telemetry vs. Survey Measurement — the instrument-weighting rule this page sharpens: randomization identifies the assigned contrast and nothing past it, so the mechanism analysis inside a randomized experiment inherits the error bar of the classifiers it runs on.
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 8
Related articles