Sources#
- Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment
- Normative boundaries of AI in scientific work: Evidence from PhD researchers
Summary#
Almost every wiki page about human involvement in AI systems argues it from the operator's side — accuracy, accountability, oversight integrity, who carries the blame. This page holds the other end: what the person judged by the system wants, and why.
Do job seekers value procedure in AI hiring only for error correction? (Chuyao Wang, Patrick Sturgis, Daniel de Kadt — LSE Methodology / LSE Data Science Institute / Cornell, arXiv 2609.16390, 2026-09-14, 32pp, empirical) answers with a preregistered paired-profile conjoint experiment: 1,919 US job seekers, eight forced choices each, 30,704 profiles, six independently randomized attributes. Its design contribution is that it varies procedure and performance orthogonally. Prior applicant-reaction work holds performance fixed, or puts the respondent outside the decision as a third-party observer — and third parties react differently from targets (Langer & Landers 2021). Here the respondent is the person the system rejects.
The two findings that matter:
- Who decides outweighs every procedural feature. Pooled human involvement is worth +0.272 in choice probability, statistically indistinguishable from the benchmark of cutting wrongful rejections from 30% to 10% (+0.285; difference 95% CI [-0.007, 0.033]).
- Procedural value does not scale with errors. Appeal, the opt-out and the bias audit are worth the same when the system rejects 10% of qualified applicants as when it rejects 30% — established by a preregistered equivalence test, not by failing to reject a null.
Together these constrain a simple error-correction account of why people want procedure. If appeal were valued for the mistakes it reverses, its value should climb with the mistake rate. It does not move at all.
The design#
| Attribute (analytic name) | Levels as shown | On-screen label |
|---|---|---|
| Decision authority | AI suggests, human decides / AI screens applicants first, human decides / AI decides alone (ref.) | "Decision process" |
| Error rate | Wrongly rejects about 10% / 20% / 30% of qualified applicants | "Accuracy" |
| Explanation | No explanation given (ref.) / Meaningful explanation | "Transparency" |
| Opt-out — human option from start | Not available (ref.) / Available | — |
| Appeal — human re-review after rejection | Not available (ref.) / Available | — |
| Independent bias audit | Not available (ref.) / Available | — |
Levels were independently and uniformly randomized subject to the two profiles differing on at least one attribute; attribute-row order was randomized per task and held fixed across the pair; task position was recorded. Estimation is OLS on the binary choice with respondent-clustered standard errors, reporting both AMCEs against the stated references and marginal means. The error rate is coded continuously in the primary model, so the registered moderation term equals the 30-versus-10 endpoint contrast — the marginal means fall in near-equal steps (0.642 / 0.499 / 0.357), which is the paper's justification for the linear coding.
What "human involvement" means here is decision authority, not review. The two human levels both leave the final call with a person; they differ in whether the AI screens first (the human never sees the rejected applications) or only suggests (the human sees everything). Neither is "a human reviews the AI's output" in the loose sense — and the paper is explicit that a supervisory human add-on is a different object, which is why its result differs from Yurrita et al.'s lending vignette where a supervisory human moved nothing.
Preregistered vs exploratory, since the distinction does real work below (pre-analysis plan posted to OSF 2026-06-17, before the data were examined; ethics approval 2026-05-28; replication code published):
- Confirmatory family — four attributes only: human involvement, appeal, opt-out, bias audit. Explanation was deliberately outside it, measured and reported separately.
- Confirmatory hypotheses: H1a (each raises choice probability), H1b (human involvement largest), H2a/H2b (does the value rise with the error rate? — tested by equivalence), H3 (appeal and opt-out gain when a human is involved), H4 (bias audit ranks below the three individual attributes).
- Exploratory / unregistered: all four feature × human interactions estimated jointly (only appeal and opt-out were registered), the "AI suggests" vs "AI screens first" × error-rate contrast, age and gender moderators, the heterogeneity simulation, and the post hoc equivalence test on the opt-out × human interaction. The paper's Appendix I is an explicit deviation ledger — an unusually honest one, since it records that a registered directional test failed and was not quietly converted into an equivalence claim.
The preference ladder#
AMCEs on choice probability, all p < 0.001 (prose §4.1 and Figure 1, which prints the same values as data labels rather than requiring a gridline read):
| Attribute | AMCE | 95% CI |
|---|---|---|
| AI suggests, human decides | +0.287 | [0.273, 0.301] |
| AI screens first, human decides | +0.257 | [0.244, 0.271] |
| (pooled human involvement) | +0.272 | [0.260, 0.284] |
| Appeal (re-review) | +0.156 | [0.145, 0.167] |
| Opt-out (human from start) | +0.129 | [0.118, 0.140] |
| Explanation | +0.128 | [0.118, 0.139] |
| Independent bias audit | +0.068 | [0.058, 0.079] |
| Error rate, per 1 pp | -0.014 | [-0.015, -0.014] |
The exchange rate is the paper's most quotable number. Twenty percentage points of wrongful rejection — the full width of the tested range — buys +0.285. Adding a human decision-maker buys +0.272. The difference is not distinguishable from zero (95% CI [-0.007, 0.033]). An employer who swaps a human decider for a more accurate machine is trading roughly one-for-one and is not obviously ahead.
Order within human involvement matters. Respondents preferred "AI suggests, human decides" over "AI screens applicants first, human decides" by 0.030 (p < 0.001) — consistent with procedural justice falling when AI is granted more decision power than its perceived ability warrants (Jiang et al. 2023). Only the first arrangement lets a human see the applications the AI would have rejected, which is the one that matters to a wrongly rejected applicant; yet in an exploratory test that margin did not move with the error rate (change across the range +0.005, 95% CI [-0.026, 0.036]). Even the part of human involvement with an obvious instrumental reading behaves non-instrumentally.
The null that carries the argument, and why it is a precise null#
A nonsignificant interaction is not evidence of negligible moderation, so H2 was preregistered as a two one-sided-tests equivalence procedure (α = 0.05, equivalently a 90% CI inside the bound) against a prespecified bound of ±0.05 on the choice-probability scale — the largest change in an attribute effect the authors were willing to call substantively negligible.
| Attribute | Moderation over the 10-30% range | 90% CI | Equivalent at ±0.05? |
|---|---|---|---|
| Appeal | +0.001 | [-0.020, 0.022] | Yes (also at ±0.03) |
| Opt-out | +0.005 | [-0.016, 0.026] | Yes (also at ±0.03) |
| Explanation | +0.010 | [-0.011, 0.030] | Yes (fails ±0.03: upper bound +0.0304) |
| Bias audit | +0.012 | [-0.009, 0.032] | Yes (fails ±0.03: upper bound +0.0323) |
| Human involvement | -0.046 | [-0.067, -0.025] | No (needs ±0.07) |
This is the shape a claim of "no effect" has to have to be worth anything. The CIs are roughly ±0.021 wide at 90% — tight enough that a moderation half the size of the bias audit's own main effect would have been detected. Appeal is the diagnostic case: of the four, its corrective function is the most direct (a rejected applicant invokes it to get human re-review), its predicted slope under error correction is the steepest, and its estimated slope is +0.001. The error-correction prediction (H2a) is supported for none of the four attributes; the performance-independence criterion (H2b) is met for appeal, the opt-out and the bias audit, but not for human involvement.
The human-involvement exception is real and the authors decline to interpret it. Its advantage over AI-alone was 0.285 at a 10% error rate, 0.289 at 20%, and 0.239 at 30% — flat, then a drop only at the highest level. The registered continuous-coding interaction (-0.046) is significant under the linear probability model (p < 0.001) and not under logit (p = 0.21). Non-monotone, scale-sensitive, and opposite in sign to the error-correction prediction — so it does not show that human involvement loses value as errors rise, and it certainly does not rescue the instrumental account. Appendix I records exactly this: "reported as scale-sensitive and not interpreted substantively."
A right you can invoke outranks oversight you cannot#
The ranking is not a smooth decay. It separates on who can invoke the feature:
- Decision authority — settles who makes the call: +0.272.
- Appeal and opt-out — routes the individual applicant can personally invoke (voice and exit, in Hirschman's terms): +0.156 and +0.129.
- Independent bias audit — system-level, group-level, diffuse, and offers no individual re-review: +0.068.
The planned contrasts are unambiguous: human involvement − appeal +0.116, − opt-out +0.143, − explanation +0.144, − bias audit +0.204 (H1b, all p < 0.001); appeal − bias audit +0.088 and opt-out − bias audit +0.061 (H4, both p < 0.001).
That maps directly onto the three regulatory instruments in force, which assign these functions differently. The EU AI Act classifies specified recruitment systems as high-risk and requires human oversight (Annex III(4)(a), Arts. 14 and 26(2)). GDPR Art. 22 grants a qualified right not to be subject to solely automated decisions, with rights to human intervention and to contest. New York City's Local Law 144 requires a bias audit plus a notice explaining how to request an alternative process — which the implementing rules do not oblige the employer to actually provide. On these estimates the three are not substitutes: LL144 legislates the attribute applicants valued least, and legislates the notice of an alternative without the alternative. Preferences do not settle legal adequacy, but "oversight alone does not exhaust what applicants value" is a measured statement here rather than an assertion.
The design's own caution runs the same way. A nominal human in the loop gives the applicant nothing to invoke either: reviewers defer to the recommendation they are meant to check (Alon-Barkat & Busuioc 2023; Green 2022) or degrade accuracy when they intervene (Sele & Chugunova 2024). The attribute that scored +0.272 was decision authority, and nothing in the design tests whether a human who nominally holds it exercises it.
Complementarity is weaker than the theory wanted#
H3 predicted appeal and the opt-out gain value when a human is already involved (voice and exit matter more when an authority can act on them). Result: partial support.
- Appeal × human involvement: +0.025, 95% CI [0.004, 0.047], p = 0.021 in the registered two-interaction model — but the same interaction is not significant under logit (p = 0.31).
- Opt-out × human involvement: -0.0003, p = 0.99. Effectively zero. A post hoc equivalence test is reported descriptively and explicitly not treated as support for H3. The design does not identify why.
In the exploratory joint model, appeal (+0.026), explanation (+0.027) and the bias audit (+0.029) each gained slightly under human involvement on the LPM scale and survive a Holm correction there — but only the bias-audit interaction survives logit (+0.117, p = 0.036). The honest summary is the paper's own: complementarities did not hold across scales, and the design cannot say which combination best protects applicants or improves decision quality. Anyone designing a feature bundle from these estimates is extrapolating.
Applying is not accepting#
A separate measured result, and the one with the widest reach outside hiring. On five-point items: intention to apply to employers using AI hiring averaged 3.24, above belief in the legitimacy of AI hiring at 2.85 (paired difference 0.389, d_z = 0.42, t(1918) = 18.56, p < 0.001). The two are strongly associated (r = 0.65) and measure different constructs — a normative judgment and a behavioral intention. 7.8% of the sample (n = 150) combined below-midpoint legitimacy with above-midpoint intention to apply (Figure 4 — the outlined block sums to 106 + 8 + 30 + 6 = 150, which is 7.82% of 1,919). Intention to withdraw averaged 2.52 and correlated negatively with both (r = -0.46 legitimacy, r = -0.67 intention to apply).
Application volume is therefore an imperfect indicator of whether applicants accept a system as legitimate — people apply to employers whose procedures they doubt, because the alternative is not applying. This is the measurement warning that travels: any behavioral proxy for acceptance collected where the subject has no outside option is measuring compliance, not consent.
Stated preference here, revealed choice there#
The sharpest tension in the vault is with Controlled Variance: AI's Edge as Reduced Dispersion, and it resolves — but only because the two studies are choosing over different objects.
Jabarian & Henkel's field experiment found 78.4% of Filipino entry-level applicants chose the AI voice interviewer when offered the choice, with offer acceptance statistically unchanged and no perceived-fairness difference between arms. This paper finds US job seekers paying +0.272 in choice probability for a human decider — an enormous stated premium.
The reconciliation is that in the field experiment humans made every hiring decision in both arms. The AI conducted the interview; it never held decision authority, which is precisely the attribute this conjoint prices. Applicants who picked a 24/7 voice agent over scheduling with a recruiter were choosing a modality under an unchanged authority structure — the arrangement this paper's respondents rate highest is "AI suggests, human decides," which is roughly what the field site ran. The two results are consistent, and the pair is a warning against reading either as "applicants do/don't like AI hiring": what applicants price is who decides, and an instrument that varies anything else will find them indifferent.
Three residual differences that the reconciliation does not dissolve, and that no design in the vault has crossed: population (US Prolific job seekers vs Philippine entry-level BPO applicants), stakes (hypothetical forced choice vs a real application), and outside option (a conjoint respondent can prefer without cost; an applicant who refuses the AI may simply not be interviewed). The legitimacy/intention split above bites hardest exactly here — the 78.4% is a behavioral statistic collected from people with a live application at stake, which is the class of measure this paper shows diverges from normative acceptance.
What the design cannot do#
Stated in the paper, and worth carrying:
- Stated choices in a hypothetical task, not application behavior. No status-quo option and no reject-both option, so every task forces a choice between two systems.
- Prolific, US-prescreened for current job seeking; n = 1,919 (soft launch 150 + 1,769). Mean age 36.6 (SD 12.1), 54.6% female, 58.6% white, 35.0% unemployed and job-seeking, 74.9% with prior experience of AI evaluation. Explicitly not representative of the wider job-seeking population. 92% passed the manipulation check; the primary analysis conditions on nothing.
- Binary attributes record whether a feature was offered, not its quality — and every feature was presented as costless, whereas real appeal and re-review involve delay, effort and uncertain outcomes. The appeal level told respondents re-review was available and said nothing about how well it worked.
- The mechanism behind the flat moderation is unidentified. The paper names two rival readings it cannot exclude: applicants may doubt that appeal works as errors mount, or may read the availability of appeal as a signal that true performance exceeds the displayed rate.
- Robustness is unusually tight. No data-quality restriction, positional control, or unordered error-rate coding moves an AMCE by more than 0.003 applied one at a time (0.005 with all three quality flags dropped at once); the five-point acceptability outcome reproduces the signs and ordering (rank correlation 0.89); soft-launch and remaining batches agree within sampling error. The fragility in this paper is in the interactions across link functions, not in the main effects.
Connections#
-
Controlled Variance: AI's Edge as Reduced Dispersion — the revealed-preference counterpart and the vault's other hiring experiment, treated at length above. 78.4% of applicants chose the AI interviewer in a design where humans held decision authority in both arms; this conjoint prices that authority at +0.272 and shows why the two do not contradict. The pair also splits the fairness evidence: the field experiment found no perceived-fairness difference between arms and halved reported gender discrimination without moving the realized offer gap, while this study finds a large stated preference on an attribute the field experiment never varied
-
Human-AI Accountability Redesign — the decision-subject's side of the same boundary question. That page's five-pillar prescription and CIVIC-AI audit both locate the human-agent boundary by verifiability, reversibility and stakes — properties of the work. This adds a constraint from outside the firm: the people judged by the system rank a right they can personally invoke (appeal +0.156, opt-out +0.129) far above system-level oversight they cannot (audit +0.068), and rank decision authority above both. A redesign that satisfies every internal audit criterion can still fail the applicant if it leaves nothing to invoke
-
Configurable Human Participation — HAS-Bench treats human participation as a configurable variable and scores configurations by task outcomes; this scores configurations by what the affected party will accept, and finds the two need not agree. Both land on the same meta-result — the configuration is what matters, not the amount of human — but from opposite objective functions, and the gap between them is the question of whose preferences a human-in-the-loop design is optimizing
-
The Automation–Optimism Link — the served-by-AI population, measured on a different construct. That page's open question asks whether optimism tracks being served by AI rather than delegating to it; here a population that has overwhelmingly been evaluated by AI (74.9%) still pays a large premium for a human decider, and prior AI-evaluation experience does not condition that premium (interaction -0.013, p = 0.36). The legitimacy/intention split (3.24 vs 2.85, 7.8% doubting-but-applying) also qualifies the behavioral evidence that page leans on
-
AI Adoption in Scientific Work — the same process-over-outcome preference, held by the people who would be using the AI rather than judged by it. Angelini & Lyrvall's 3,785 PhD students accept AI for literature work (51.5% comfortable with summarising) and refuse it for writing, data analysis and experiment design (~30%) — a boundary drawn around which step the machine performed, independent of whether the output would be any good, which is this page's central claim transplanted into scientific evaluation. Their §5.1 states it directly: a governance regime "concerned exclusively with output quality may therefore overlook differences in authorship, accountability, and the distribution of intellectual contribution across the research process." Two caveats keep the pairing honest — that instrument is unrandomized comfort, not a causal AMCE, and it measures no willingness to trade, so it prices nothing
-
Telemetry vs. Survey Measurement — a fourth instrument for that page's ledger: a randomized stated-preference design. It is not telemetry (nothing is observed) and not a survey in the attitudinal sense (attributes are randomized, so AMCEs are causal within the vignette), and it buys the one thing neither of the others can — orthogonal variation in procedure and performance. It pays for that with external validity, and this paper's own legitimacy/intention gap is a direct measurement of how far stated acceptance and behavior can separate
-
The Enablement–Regulation Axis — the supply side of the same question, and it does not answer the demand this page measures. Applicants rank individually-invocable procedure highest: human decision authority +0.272, appeal +0.156, opt-out +0.129, against a system-level bias audit at +0.068. Chueri & Törnberg's 33-parliament corpus finds the legislative topic containing exactly those instruments — algorithmic management, platform-work safeguards and human oversight — is the smallest component of the regulation-and-restriction frame at 13.0% of it, or 155 mentions out of 5,470 total response-frame mentions, while copyright and creative-sector protection leads that frame at 36.1%. Compensation, the post-hoc remedy this paper's respondents were never offered, is 2.3% of all mentions. The two studies observe different populations, countries and periods and neither establishes a representation gap alone — but they are the vault's only measurements of the two ends of one, and they point the same way
Into hubs (one-way)#
- Verification as the New Bottleneck — the applicant-side reason the bottleneck cannot simply be automated away. A nominal reviewer who defers to the recommendation satisfies an oversight requirement while giving the decision subject nothing to invoke, and the attribute applicants actually priced was who holds authority, not who is nominally in the loop
Open Questions#
- Every procedural feature in this design was costless: the level said "available" and never said what invoking it costs in delay, effort or uncertainty. Does the flat moderation survive when appeal is priced — e.g. randomizing "re-review within 3 days" against "re-review within 6 weeks," or attaching an explicit success rate to it? An error-correction account should revive as soon as the feature's expected corrective yield is made legible, and the equivalence bound is tight enough (±0.021 at 90%) that a real slope would show.
- The paper names two rival explanations for the flat moderation it cannot separate — applicants value procedure for standing and dignity, or applicants discount appeal's efficacy exactly as errors mount (and/or read its availability as a signal that true performance beats the displayed rate). These predict identical AMCEs and opposite policy conclusions. Does a design that measures believed appeal-efficacy alongside choice, or that states the appeal success rate as its own randomized attribute, split them?
- The vault now holds a stated premium for human decision authority (+0.272, US Prolific job seekers) and a revealed 78.4% choice of an AI interviewer (Philippine entry-level applicants, humans deciding in both arms). The object-of-choice reconciliation above accounts for the sign, but not the magnitude: is the residual attributable to population and stakes, or would US job seekers facing a real application also take the AI? Answerable by synthesis over the two pages plus the vault's other stated-vs-revealed pairs before any new source arrives.
Sources#
- Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment — Chuyao Wang, Patrick Sturgis (LSE Department of Methodology; Wang also LSE Data Science Institute) & Daniel de Kadt (Cornell), Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment, arXiv 2609.16390 (2026-09-14; 32pp, 18 tables, 6 figures;
empirical). §3.1 sample and ethics (Prolific, US residence + current job seeking prescreen, n = 1,919, mean £10.22/hour, 92% manipulation-check pass); §3.2 the six attributes and their on-screen wording; §3.3 estimator, AMCE/marginal-mean definitions and the preregistered ±0.05 TOST; §4.1 the AMCE ladder and the 20-pp error-reduction benchmark; §4.2 the equivalence results and the human-involvement exception; §4.3 the complementarity interactions; §4.4 legitimacy vs intention to apply; §4.5 robustness; §5 the regulatory mapping (AI Act Arts. 14/26(2) and Annex III(4)(a), GDPR Art. 22, NYC Local Law 144) and the stated limitations. Appendices: B (sample characteristics), C1-C3 (AMCEs, planned contrasts, marginal means), D1 (equivalence tests), E1-E2 (subgroup and full interaction estimates, LPM and logit), F (robustness), G (power — target ~1,400 set by the moderation and opt-out-by-human tests, Monte Carlo on a separate excluded pilot), H (heterogeneity), I (deviations from the pre-analysis plan). Preregistration: OSF https://osf.io/5ju4d/ posted 2026-06-17 before the data were examined; LSE Methodology ethics approval 2026-05-28; replication data and code published. Declaration of no competing interests; the funding statement is an unfilled template placeholder in this preprint, so funding is undisclosed rather than declared absent. - Figures viewed and reconciled. Figure 1 (AMCE forest plot) prints its estimates as explicit data labels that match Appendix C1 exactly (+0.287 / +0.257 / +0.156 / +0.129 / +0.128 / +0.068, error rate per -1 pp +0.014) with the 20-pp error-reduction reference line at +0.285. Figure 2 Panel B reproduces the prose's human-contrast series (0.285 / 0.289 / 0.239) and Panel A the equivalence intervals. Figure 3's printed deltas match Appendix E2 Panel B (bias audit +0.029, explanation +0.027, appeal +0.025, opt-out -0.000). Figure 4's outlined block sums to 150 respondents (106 + 8 + 30 + 6), confirming the 7.8% in the prose. Figure 5 confirms the ±0.03 / ±0.05 / ±0.07 bound sensitivity. Per vault convention no number on this page is read off a chart — every figure quoted comes from prose or from a table reconciled against
pdftotext -layout. - Parse notes — every cited table reconciled, two benign docling artifacts, one unverifiable table. The raw carries no
[!note]/[!warning]repair blocks, so all 18 tables were treated as unrepaired and checked cell-by-cell againstpdftotext -layouton the local PDF. Tables 1, B1, B2, C1, C2, C3, D1, E1, E2 and I1 are byte-exact, including signs, CI brackets and the minus-sign glyphs. The ingest verify's fourtable-collapseand twotable-weldflags are confirmed false positives: they fire on multi-level summary rows ("Error rate: 10% / 20% / 30%", "Prior AI evaluation: yes / no / not sure") and on two-line wrapped labels ("Decision authority: AI screens first, human decides", "Prior AI experience: no or not sure"). Two artifacts to know about: (i) Tables C1, C3, E1 and B1 each appear in the raw as two markdown tables, because the PDF splits them across a page boundary and repeats the column header — no rows are lost, and the split points are 24/25, 25/26 and 26/27; (ii) the tail of Table B1 (Mixed, Other, employment status, prior AI evaluation, country) renders in the raw below the "Table B2: Attribute-level frequencies" caption, so the caption appears to introduce demographic rows it does not describe — the PDF order is continuation-then-caption, and B2 proper is the attribute-share table beneath it. Table 2 (the example paired-profile task) is a screenshot of the instrument with no PDF text layer, so it exists in the raw only as RapidOCR output and cannot be reconciled againstpdftotextat all; the ingest pass repaired oneAI→Almisread in it. Nothing on this page is cited from Table 2 — the design description comes from Table 1 and §3.2 prose, both verified. - Normative boundaries of AI in scientific work: Evidence from PhD researchers — Angelini & Lyrvall, Normative boundaries of AI in scientific work: Evidence from PhD researchers, arXiv 2608.25678, 2026-08-26, 15pp,
empirical. Cited here only for the process-versus-outcome governance argument in §5.1 and the task-level comfort marginals in Table 1, as the practitioner-side analogue of this page's decision-subject preferences. Full treatment on AI Adoption in Scientific Work
Cited by 10
- AI Adoption in Scientific Work
Procedural Value In Ai Decisions — the same process-over-outcome preference measured on the people…
- The Automation–Optimism Link
Procedural Value In Ai Decisions — the served-by-AI population measured directly, on a different…
- Configurable Human Participation
Procedural Value In Ai Decisions — the same configuration sweep scored on a different objective.…
- Controlled Variance: AI's Edge as Reduced Dispersion
Procedural Value In Ai Decisions — the stated-preference counterpart, and the page that explains…
- The Enablement–Regulation Axis
Procedural Value In Ai Decisions — the two halves of one representation gap, measured independently…
- Human-AI Accountability Redesign
Procedural Value In Ai Decisions — the constraint from outside the firm. Every pillar here locates…
- AI Economics & Labor
Procedural Value In Ai Decisions — Wang, Sturgis & de Kadt (LSE/Cornell, arXiv 2609.16390): a…
- Open Questions Dashboard
Procedural Value In Ai Decisions: The vault now holds a stated premium for human decision authority…
- Open Questions Backlog
Procedural Value In Ai Decisions ×3 (oldest 6d) — Every procedural feature in this design was…
- Telemetry vs. Survey Measurement
Procedural Value In Ai Decisions — an eighth aperture: the randomized stated-preference design. A…
Related articles
- Organizational Complements to AI
The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…
- Controlled Variance: AI's Edge as Reduced Dispersion
Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants betwe…
- The Tragedy of the Cognitive Commons
Lovett (HRD Review, July 2026): professional expertise is a profession-level commons whose regeneration mechanism — ent…
- AI Brain Fry
Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognit…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated
Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoreti…
