H
Howardism
Plate IIAI Coding Practice中文HOWARDISM

Telemetry vs. Survey Measurement

PublishedJune 17, 2026FiledConceptDomainAI Coding PracticeTagsEngineering MetricsMeasurementAI Coding WorkflowReading34 minSourceAI-synthesised

Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (instrument agreement tracks shared data source, not construct), randomization as the only causal instrument, the survey arm's counter-case (shadow AI is invisible to telemetry), and the Ramp payment-rail aperture; first cross-family convergence: Anthropic passing OpenAI mid-2026.

Illustration for Telemetry vs. Survey Measurement

Sources#

Summary#

Faros AI's methodological argument, and the basis for the most consequential conflict in its 2026 report: during rapid AI transformation, perception lags reality, so survey-based engineering research systematically misses the downstream damage that system telemetry catches in near-real-time. Faros draws its findings from engineering systems (task trackers, IDEs, static analysis, CI/CD, version control, incident management) rather than from how developers feel, and uses that distinction to directly contradict Google's DORA 2025 conclusions.

vendor-claim source — Faros's own platform is the telemetry instrument, so "telemetry beats surveys" is also a sales argument for that platform. The methodological point stands on its own merits, but the conclusion conveniently favors the vendor's product. See Acceleration Whiplash for the full evidence note.

Why perception lags reality#

The mechanism Faros proposes: at the individual level developers genuinely are more productive — task completion is up, code flows faster, the tools feel powerful — so surveys capture real, positive feeling. What surveys cannot capture is what happens downstream: "the review queues quietly backing up, the incidents accumulating in production, the bugs reaching customers." By the time those consequences show up in how people feel, "months have passed and the signal is already stale." Telemetry, drawn from the systems where work actually happens, does not lag. The claim: engineering leaders making consequential decisions about headcount, tooling, and process "need data as close to real time as possible… not how people feel about the work after the fact."

The DORA contradiction#

This is a flagged inter-source contradiction. DORA's 2025 State of AI-Assisted Software Development concluded that AI amplifies existing strengths and weaknesses, and that strong engineering foundations protect against AI's downsides. Faros's telemetry, it claims, "does not support that as a protective factor": high-performing organizations experience the same downstream deterioration as everyone else (see the maturity-independence finding in Acceleration Whiplash).

Weighing the conflict by method and incentive:

  • DORA 2025 — survey-based; large, long-running, vendor-neutral-ish (Google/DevOps Research). Strength: breadth and continuity. Weakness, per Faros: perception lag during fast transitions.
  • Faros 2026 — telemetry-based; within-company longitudinal comparison (low- vs high-adoption quarters), Spearman ρ at p<0.05. Strength: measures behavior, not feeling, near-real-time. Weakness: vendor-claim — Faros sells the platform, and "your mature practices won't save you, you need visibility + a context engine" is precisely the conclusion that grows its market.

Neither is a clean win. The honest read: Faros's measurement critique of surveys is sound (lagging perception is real), but its substantive claim that maturity offers zero protection should be held with the vendor incentive in view — it is the conclusion most favorable to selling the instrument. Worth tracking against future DORA editions and any non-vendor telemetry study.

The family effect: instruments agree with their data source, not with their construct#

The sharpest evidence that this page's dichotomy is a real fault line rather than a framing device comes from outside engineering metrics. Steele & Cruz (arXiv 2607.15506) put seven occupational AI-exposure instruments on the same O*NET occupations and correlate them pairwise. All seven claim to measure the same thing. What predicts whether two of them agree is which data source they were built from:

  • The two built from 2025 Anthropic Claude usage correlate at ρ = 0.89.
  • The two built from theoretical generative-AI capability (GPT-4 task ratings; 2,000 MTurk ability ratings) are the next-strongest pair — despite differing in level of analysis, rater, and question asked.
  • Across families, correspondence largely collapses; patent-mining, ML-rubric, and bottleneck-based instruments show "very little correspondence with each other or with later measures."

Two consequences for this page. First, the telemetry/survey split is not just a latency difference (behavior now vs. feeling later) — it is a partition of the answer space. Instruments in different families do not produce the same ranking, and in Steele & Cruz's case they do not even produce the same sign on the exposure-salary gradient. Second, the paper is the first entry in the vault where telemetry is run forward into a projection rather than reported as an observed count: query-volume ventiles rank the tasks, and a hand-set schedule of assumed automation ceilings supplies the levels. That is a hybrid — telemetry's ordering, assumption's magnitudes — and it inherits the credibility of the ranking without inheriting it for the numbers. See Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated for the full head-to-head.

The third instrument: randomization, and the one thing neither telemetry nor survey can do#

This page's dichotomy has a missing third term. Telemetry and surveys are both observational — one records behavior as it happened, the other records how it felt. Neither can say what would have happened otherwise, which is the question every claim on this page is ultimately about. A randomized experiment can, and Jabarian & Henkel (arXiv 2607.28222) is the vault's cleanest instance: 67,056 job applicants randomized between AI voice interviewers and human recruiters, pre-registered, with administrative employment outcomes.

What makes it useful here is that the paper runs all three instruments on the same event and they point in different directions:

InstrumentWhat it says about AI-conducted interviews
Administrative records (offers, starts, retention)AI-interviewed applicants receive 12% more offers, start more often, and are retained more (all p<0.001)
Applicant survey (post-interview CX survey)AI interviews are rated significantly less natural (p=0.014), pulling the composite perceived-quality index in the humans' favor (p=0.076)
Expert forecast (recruiter survey, pre-disclosure)61% of recruiters expected AI-led interviews to be of lower quality; 36% expected lower offer rates; 48% expected lower retention

The forecast is simply wrong — 131 experienced practitioners, surveyed before results were disclosed, predicted the sign backwards. The applicant survey is not wrong; it measures a real and different thing (naturalness) that does not track the outcome. And the administrative records are the only reason we know which is which.

Two consequences for this page. First, the felt-vs-system split it was built on is a special case of a bigger problem: an observational instrument is a claim about correlation regardless of how good its latency is, which is why CMU's 2.5M-PR telemetry study could support opposite conclusions from the same traces. Randomization is the rung above both, not a third flavor at the same level. Second, the honest cost of that rung: it took a firm partnership, an IRB, a pre-registration, and three months of a live hiring pipeline to answer one question about one task in one labor market. Telemetry and surveys are cheap and broad; the instrument that actually identifies the effect is expensive and narrow, which is exactly why the vault has hundreds of the former and one of the latter.

The survey arm, made concrete: what only asking can see, and what it costs#

Everything above argues against surveys. Kalff & Simbeck's N=410 study of German HR departments (arXiv 2607.13839, empirical) is the survey arm of this contrast in its plainest form — a self-reported adoption census of one business function, whose authors list "reliance on self-reported survey data for quantitative measures" among their own limitations. It is useful here for the opposite reason to the rest of the page: it contains the clearest instance the vault has of something no telemetry can see, and the clearest instance of the price of asking.

The thing only a survey can measure. Of the 410 valid respondents, 183 report using AI tools informally, on personal devices, regardless of employer policy or the availability of company-provided systems — including at companies whose official policy restricts or prohibits generative AI. The paper draws the boundary itself: "AI solutions requiring complex, company-specific data cannot be used informally, as they are accessible only through an official rollout." So corporate telemetry sees exactly the half of adoption that went through IT, and is blind by construction to a channel roughly as large. This is the CircleCI aperture problem in its limiting case — the instrument's coverage is whatever the vendor's product touches, and shadow adoption touches nothing the firm operates. On this question the ordering inverts: the survey is not the lagging instrument, it is the only instrument.

And the price of asking, in an unusually legible form. The same paper documents its own construct collapsing. Group discussions "often devolved into debates over the meaning of AI"; participants' perceptions were "strongly shaped by their exposure to generative AI applications popular in the media," so "some participants did not recognize machine-learning systems that underpin traditional HR analytics — such as those used to measure turnover risk or identify workforce trends — as 'actual AI.'" That is not ordinary measurement noise, because it is directional and it points at the paper's own headline: the conclusion is that predictive analytics plays only a minor role, and the tools respondents are least likely to count as AI are precisely the predictive-ML ones. A license- or spend-based instrument (Ramp's payment traces) has no such problem — it counts what was deployed, not what gets called AI.

Worse, the ambiguity is not merely fuzzy, it is contested by interested parties in both directions: vendors "leverage AI branding to generate interest and support business cases," while "the AI aspects of HR tools may be downplayed to avoid scrutiny from co-determination bodies." When the label is a strategic instrument for the people answering the question, no amount of question wording recovers the construct — a third failure mode, distinct from perception lag (this page's original argument) and from the family effect (instruments agreeing with their data source). Perception lag says respondents report an old truth; the family effect says instruments disagree about a stable construct; this says the construct itself is being moved by the respondents while you measure it.

The payment rail: the platform-records family gets a second member, and the first case where survey and behaviour agree#

The Connections entry below names Indeed's job postings as a fourth instrument family — a platform's by-product record of transactions it brokers, neither telemetry nor survey nor randomization. Ramp's monthly AI Index (empirical) is that family's second member, and putting the two side by side sharpens what the family is: whoever runs the rail can count what crosses it, for free, in near-real time, for exactly the population that transacts there. Indeed brokers vacancies and can therefore count labor demand; Ramp brokers corporate payments and can therefore count who firms pay for AI. Both are published by the operator's own research arm, so both fold the vendor incentive and the instrument into one party the way this page's Faros entry does.

The aperture bites in a checkable way here. Ramp's June-2026 vendor shares put OpenAI at 39.5% and Anthropic at 42.4% of businesses — and Google at 6.4% and Microsoft at 1.7%, with Google flat in the 4.6–6.4% band for the entire 42-month series while overall AI adoption went 7.5% → 55.0%. The most likely reading is not that Microsoft and Google lost the enterprise; it is that enterprise agreements, negotiated invoices and cloud committed-spend drawdowns do not cross a corporate-card rail the way a per-seat or per-API subscription does. That is the CircleCI aperture problem restated for payments: the instrument's universe is the purchases that fit its rail, and a vendor whose sales motion routes around that rail is structurally undercounted regardless of its actual share.

And the part this page did not previously have an example of: a survey and a behavioural instrument agreeing. The family effect above predicts that instruments track their data source rather than their construct, and most of this page is disagreements. Here two instruments from different families, on different populations, converge over the same six months on the same reordering: ICONIQ's Q2-2026 exec survey of ~305 AI-building software companies has Anthropic going 51% → 81% of respondents and passing OpenAI (77% → 71%), while Ramp's payment records have Anthropic passing OpenAI in May 2026 (42.4% vs 39.5% by June) with OpenAI down ~1.9pp from its November-2025 peak. Self-report and receipts, a builder cohort and a whole card base, same direction and roughly the same timing. Convergence across families is weak evidence taken alone and strong evidence taken here, precisely because this page's default expectation is that it does not happen — and it is the cleanest case in the vault of the instrument question being settled by agreement rather than adjudicated.

"The control group is no longer viable" — a claim to split, not to accept (2026-08)#

DX's Q2 2026 AI-impact readout (vendor-claim, 500+ customer organizations) opens with the sharpest methodological assertion any source in this vault has made about its own instrument, and it is aimed squarely at this page's subject:

"When we first began tracking the impact of AI on engineering teams, our primary goal was to measure AI cohorts against historical baselines… With industry-wide AI adoption exceeding 90%, comparing AI users against a non-user control group is no longer a viable measurement strategy."

The claim is true of one control design and false of another, and the two are not usually distinguished. What saturates at >90% is organization- and developer-level adoption — whether a person or a firm uses AI at all. What has not saturated is the artifact-level mix: the share of individual changes, PRs or completions that an AI actually wrote.

  • The between-firms (or between-developers) adopter-vs-non-adopter contrast is genuinely dying, and DX is right about it. At >90% adoption the non-adopter arm is a residual of holdouts selected on whatever made them hold out, which is a worse confound every quarter. This is also the design DX's own panel and Faros's adoption-depth cross-section are built on, and the reason neither can separate "AI code is worse" from "orgs that adopt hardest differ in other ways."
  • The within-firm, within-codebase provenance contrast is alive and was running while DX declared it dead. Tran et al. (Google) compare AI-authored against human-authored changes inside one monorepo, same window, stratified on change size, over 3.52M submissions. Adoption in that population is effectively total — everyone has the tools — and the control cohort survives anyway, because AI's share of submitted code ran 28.99% to 68.62%, leaving roughly three in ten changes human-written at the end of the window. Universal adoption and a usable control cohort coexist without tension the moment the unit of comparison stops being the person.

So the correct statement is narrower and more useful than DX's: the control group moved down a level rather than disappearing. The population-level arm is gone; the per-change arm is not.

Why the claim is also structurally self-serving, in the way this page's Faros entry already documents. DX sells developer-productivity measurement, and its instrument is a survey-plus-SDLC-telemetry panel across customer organizations — an instrument that can only ever run the between-firms contrast. The design it declares non-viable is the one it cannot run; the remedy it proposes (longitudinal within-panel trends against historical baselines) is the product. That does not make the observation wrong, and the honest reading is that DX has correctly identified a real measurement crisis and misattributed its cause.

The cause is instrumentation, not adoption. A per-change control cohort requires authoring-time provenance — a record, made while the code is being written, of what wrote it. That is a decision taken before the artifact exists, by whoever owns the editor, and it is why Google could build both arms and nobody outside such a company has. The alternative available externally is a corpus defined by agent authorship (AIDev), which yields one arm by construction and no matched human comparison — which is exactly why Dipongkor et al. and Jhanglani et al. and Sakib et al. each measure a level and not a delta, and why DECODE can see only accepted completions. Four studies missing the same arm, for a reason that has nothing to do with adoption rates. Post-hoc AI-code detectors are the obvious substitute and Tran et al. reject them outright ("generalize poorly across models and settings").

The prescription that follows for reading this vault: when a source says it could not build a control group, ask whether it lacked non-adopters (increasingly unavoidable) or lacked provenance (a fixable instrumentation gap that most publishers of AI-impact numbers have not paid for).

Connections#

  • Community Smells Under AI Adoption — the instrument tension appearing inside one instrument: the same survey's free-text layer reports AI displacing teammates while its structural models estimate peer interaction rising. The authors' three-part reconciliation — individual change vs cross-respondent association, concentrated-and-conditional effects, and a non-significant direct path indicating mediation rather than absence — is a reusable frame for reading self-report here

  • The Open-Weight Frontier Gap — the limit of a payment-rail instrument on the question it is most often used for: Ramp's 5.8%-of-AI-spenders model-serving figure is the vault's main quantitative handle on open/Chinese-model adoption, and it can only see firms that pay a serving vendor — downloading open weights and running them on your own GPUs generates no transaction and is invisible by construction, the same blind spot shadow AI creates for corporate telemetry above

  • Organizational Complements to AI — the substantive home of this source, and the reason its instrument matters: adoption below the organizational level (183 of 410 using AI outside any rollout) is a complements story that firm-level telemetry cannot reach, so the complements literature's usage-share and spend-trace instruments are structurally blind to the fastest-diffusing half

  • Controlled Variance: AI's Edge as Reduced Dispersion — the third instrument above: the vault's only randomized causal estimate of AI substituting for a human in an expert conversational task, and a single design in which administrative records, an applicant survey, and an expert forecast disagree about the same intervention

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — where the seven-instrument comparison lives; the family effect above is the measurement-methodology reading of it, and its telemetry-built seventh instrument is the hybrid case this page's dichotomy did not previously have a slot for

  • Usage-Telemetry Classifier Validation — the price of the telemetry side of this argument: system logs beat self-report on latency and scale, but the economic claims drawn from them run through an LLM classifier whose exact accuracy (22.6% at the O*NET task level) was never published until Google ATLAS did it

  • AI Native Product Cadence — the limit of the telemetry side: OpenAI's Akshay Nathan reports the classic engineering proxies (commits, LOC, PRs, tokens) decoupling from goal attainment under agents, so better instrumentation of the same signals measures motion rather than progress — telemetry beats survey on latency, but only if the logged quantity still tracks the outcome

  • Acceleration Whiplash — the maturity-independence finding rests on this telemetry-over-survey methodology

  • Production-Sourced Evaluation — the same "measure from the real system, not a proxy" instinct applied to model evals; telemetry-vs-survey is its engineering-metrics cousin

  • Evals as Product SpecCat Wu's evals encode the spec; telemetry encodes what actually shipped — both prefer ground-truth signal over self-report

  • Verification as the New BottleneckFiona Fung's warning to break PR-cycle-time into funnel chunks rather than read the aggregate is the same "instrument carefully or the signal misleads" discipline. That page also carries the instrument-aperture case: CircleCI's Q2 2026 Pulse (vendor-claim, 20M+ workflows) is behavior-not-feeling telemetry like Faros's, but from a single platform on a single signal — it sees CI workflow runs and nothing else. That aperture fixes what it can claim: it counts validation cycles (its Merge Efficiency Ratio) with a precision no survey could reach, and cannot observe a production incident, a customer-facing bug, or a review queue at all. So its cheerful main-branch success rate (70.8%→76.7%) and Faros's grim incidents-per-PR (+242.7%) are barely in conflict — each vendor reports the half of the SDLC its own product instruments. The corollary for this page's thesis: telemetry beats surveys on latency, but its coverage is whatever the vendor's product happens to touch, and a vendor's headline metric is reliably one its instrument can see and its product can improve

  • Compounding Data Moat — owning the telemetry stream is itself a moat; the report is a demonstration of what the data asset enables

  • Agentic Coding Work-Composition Shift — the empirical cousin: Anthropic's Clio-based 400K-session telemetry reads behavior-not-feeling the same way, but is a research artifact (validated classifiers, controls) rather than a vendor-claim lead-gen report — and its session-layer optimism vs Faros's org-layer pessimism is the felt-vs-system split this page names

  • Conversation-to-Delegation Shift — the third major usage-telemetry study (OpenAI/Codex, empirical), and it extends this page's argument: as usage becomes delegation, even interaction-count metrics (active users, chats) go stale — track complexity, runtime, concurrency, reuse, output instead

  • Anthropic Economic Index — the program that resolves this page's dichotomy: it links usage telemetry to survey responses per person (the Cadences report, ~9,700 linked respondents), treating telemetry and survey as complements rather than rivals

  • AI Usage Cadences — the AEI's continuous hourly telemetry is this page's "measure the real system finely" principle pushed to time resolution

  • Review as the Control Point — the non-vendor telemetry this page's open question asked for, and a deepening of its methodology argument: CMU re-scraped 2.5M+ GitHub PRs (behavior, not feeling) yet found the same traces support opposite conclusions under defensible analysis choices — so telemetry beats surveys on latency, but surface telemetry can't adjudicate why without a causal model (Pearl: "data are profoundly dumb"). Telemetry's edge is real and bounded

  • Market-Priced AI Exposure (the AI Premium) — the purest realized telemetry (every observation is a paid request, behavior not feeling), pushed to cross-provider breadth (400+ models, ~2% of global tokens) and then turned into a market-priced signal via equity-price comovement; the finance-side entry in this thread — but a reminder that telemetry's own collection mechanism biases it (OpenRouter is developer-skewed), so realized ≠ representative

  • AI Investment Story, Not Efficiency Story — this page's dichotomy carried into startup economics: on revenue-per-employee, Emergence's measured cap-table receipts (AI companies below non-AI peers) conflict with AWS's self-reported founder survey (AI-natives above the baseline) — a survey-vs-financial-telemetry instrument split, same felt-vs-system shape as the Faros-vs-DORA case, resolved the same way (weight the receipts, flag rather than average)

  • Firm AI-Spend Intensity and Headcount Growth — also the home of a fourth instrument family this page had no slot for: platform administrative records, in Indeed Hiring Lab's job-postings series. Not telemetry (nobody's usage is logged), not survey (nobody is asked), not randomization (nothing is assigned) — it is the by-product record of transactions a platform brokers, which makes it behavior-not-feeling and near-real-time on the demand side, where usage telemetry sees only the supply side. It inherits the aperture problem in a sharper form than CircleCI's: the instrument is one job board, so its universe is whoever posts there, and "US job postings fell 7%" is a statement about Indeed's marketplace before it is a statement about the economy. It also collapses the vendor-incentive and the instrument into one party in the way this page's Faros entry does — Indeed's research arm publishing a rebound story about Indeed's own board — while the substantive limit is different from Faros's: the numbers are administrative and hard to dispute, and it is the causal reading ("Claude Code launched, then postings rebounded") that the design cannot carry. Beyond that, the same behavior-not-feeling instinct applied to AI adoption and its labor effect: Ramp reads adoption off actual AI-vendor payment traces (not "do you use AI?" surveys, whose answers for the same period span 18%→78%) and joins them to Revelio workforce records — a revealed-adoption instrument the paper positions explicitly against "messy surveys and exposure measurements"

  • AI Product Economics Maturation — the survey axis pushed to its softest form: ICONIQ's exec survey layers forward projections on top of self-report (2026P/2027P margin and RPE), so its rosy trajectory is felt-about-the-future, two removes from telemetry. Where its projected margin expansion (→59% by 2027) meets Emergence's measured growth-margin compression, this page's prescription applies unchanged — weight the receipts, flag rather than average, and mark the projection prediction-grade

  • Outsource Your Thinking, Not Your Understanding — the reverse case to this page's thesis: comprehension debt is damage the telemetry misses, because the shipped artifact is fine (Shopify reports AI-assisted reversion rates at pre-AI baselines) and the human is what changed

  • Standardize the Infrastructure, Not the Tools — the instrument that makes org-wide AI telemetry possible at all (one gateway, one meter), and the open question of what per-team usage analytics get used for once they exist

  • Efficiency Debt of AI-Generated Codefirst-party telemetry with a control cohort, which is the shape this page's whole Faros argument has been missing. Google instruments its own monorepo at byte-level authoring provenance and compares AI-authored against human-authored code inside the same window and size stratum — so unlike Faros's low-adoption-vs-high-adoption cross-section it can separate "AI code is worse" from "orgs that adopt hardest differ in other ways." Two lessons for this page. The aperture rule holds and cuts deeper than usual: the instrument sees everything inside one company's pipeline and nothing outside it, which is total coverage of a population of one. And the vendor incentive runs the opposite direction from every other entry here — the org measuring the code is the org that built the tools that wrote it, and the result is on balance reassuring (revert rate below parity) with a discussion attributing the residual weaknesses to "historical default system configurations." Two telemetry sources, two directional COIs, opposite signs, and only one of them has a control group

  • AI and Market Power — the fifth instrument family, and the only one with a real claim to representativeness: compulsory national statistical surveys run by statistical offices to common Eurostat guidelines (France's TIC, Portugal's IUTICE), linked to administrative balance sheets. It is a survey, so it inherits the self-report limits this page catalogues — but it is a sampled, weighted, mandatory one rather than a self-selected panel, and it anchors the low end of the vault's adoption spread: 20.2% of OECD enterprises using AI in 2025, next to Census BTOS's 18%, Ramp's ~55% of eligible businesses on the payment rail, and executive surveys at 69–78%. The aperture argument gets its cleanest illustration here — five instruments, one phenomenon, a 4× spread, and the most representative one reads lowest

  • Post-Acceptance Edit Behavior — a sixth instrument family, and the only one that observes an artifact which is then destroyed: pre-commit editor telemetry. DECODE saves the working file every time a developer pauses for one second, so it sees the AI code that gets deleted 23 minutes after acceptance — code that exists in no repository, no PR, no survey response and no monorepo history, and is therefore invisible to every other instrument on this page. That is the aperture argument's mirror image: the layers below the commit are not a smaller view of the same population, they contain events the higher layers structurally cannot record. The weakness is the familiar one in a sharp form — the aperture is one opt-in VS Code extension whose users installed a model-comparison tool, self-selected in a way a mandatory statistical survey or a company's full monorepo history is not

Derived#

Open Questions#

  • Is there a non-vendor telemetry dataset large enough to adjudicate the maturity-protection question independently of Faros's commercial framing? Partially answered: CMU's arXiv 2607.07980 supplies exactly this — a non-vendor, 2.5M+-PR GitHub telemetry study — and it (a) finds the agent no-review rate converging toward the human baseline rather than a widening gap, and (b) argues the effect's sign is set by team practice, closer to DORA. The catch: its own headline is that the telemetry is direction-unstable, so it counterweights Faros without cleanly settling the maturity question — the honest verdict is "surface telemetry alone, vendor or not, can't adjudicate this." The worked example of that verdict is The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence: the two telemetry studies' opposite under-review headlines dissolve once metric, population, time axis, and authorship unit are aligned — the datasets never disagreed, only the framings did.

  • Does anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction? Kalff & Simbeck document respondents excluding predictive ML from the category while counting ChatGPT in it, which would bias reported adoption down for exactly the tool class their headline finding is about. Falsifiable two ways: a split-ballot survey where one arm names the underlying technique ("software that scores turnover risk") instead of the label, or validating firm-level self-reported adoption against vendor-license/spend records for the same firms — the Ramp instrument applied as a criterion rather than a substitute. Related evidence (2026-08-12), not an answer: DX reports a self-reported AI-generated code share of 34% to 52% across Q1-Q2 2026 in 500+ organizations, while Google's authoring-time provenance over the overlapping window reads 28.99% to 68.62%. Different populations (a cross-industry customer panel vs one C++ monorepo), different constructs (code share vs adoption), no matched firms, and DX's question wording sits in a gated PDF the vault does not hold — so this settles nothing. It is nonetheless the corpus's first side-by-side of a self-reported code share against a provenance-measured one, and the direction is the one Kalff & Simbeck's construct-collapse mechanism predicts: self-report reads below the instrument that counts bytes, not above it. The criterion validation this bullet asks for is that same comparison run on the same firms.

Resolved Questions#

  • Surveys and telemetry measure different things (felt productivity vs. system outcomes); is the "contradiction" partly a category error — both true at their own layer — rather than one being wrong? Answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — yes for the measurement halves: surveys capture felt productivity (genuinely real at the individual layer — task completion did rise) while telemetry captures system outcomes that haven't propagated into feeling yet; during a fast transition the layers legitimately diverge, and the AEI's linked telemetry+survey design already treats them as complements. But not entirely: the maturity-protection claim is one proposition about one layer (system outcomes) and remains substantively contested — Faros's "no protection" aligns with its vendor incentive, CMU's moderator theory argues the sign is team-set (closer to DORA) but measures no maturity effect. The category-error dissolution cleans the framing; the one real disagreement stays open (tracked in the non-vendor-dataset question above).

Sources#

  • AI Engineering Report 2026: The Acceleration Whiplash — "A direct counterpoint to DORA's 2025 findings"; Research Methodology; Report's Purpose
  • DORA, 2025 State of AI-Assisted Software Development (cited by Faros): https://dora.dev/research/2025/dora-report/
  • 5 takeaways from the State of Software Delivery Q2 Pulse report — Jacob Schmitt, CircleCI blog (2026-07-08), vendor-claim: the single-platform telemetry aperture and the Software Delivery Data Explorer benchmark framing
  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, arXiv 2607.28222 (2026-07-30), empirical, pre-registered field experiment. §3.1 (administrative outcomes), §3 intro (the recruiter forecast: 36% lower offers / 48% lower retention / 61% lower interview quality), §5.2 (the applicant survey's naturalness result and the composite quality index). Full treatment and parse warnings at Controlled Variance: AI's Edge as Reduced Dispersion.
  • Helping People Choose Careers in the Age of AI — Steele & Cruz, arXiv 2607.15506 (2026-07-16), empirical. §4 (the query-based construction) and §4.4 + Fig. 6 (correspondence among the seven instruments). Table 1 and Table 2 verified clean at ingest; Table 5 and appendix Table A1 are damaged in the parse and are not cited here. The paper contains a source-internal contradiction in its own summary statistics — see the Sources note on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated.
  • AI and Job Postings: From Destruction to Creation? — Guillermo Gallacher, AI and Job Postings: From Destruction to Creation? (Indeed Hiring Lab, 2026-07-08; empirical). Cited here for the instrument rather than the substance: a platform's own administrative transaction records (job postings), analyzed and published by that platform, with an uncontrolled before/after design around a product launch. Substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth; chart-only quantities in the source (per-country shares, scatter-plot correlation statistics) are not quoted anywhere in this vault.
  • Ramp's latest data on China vs. the American AI Labs — Ara Kharazian, Ramp's latest data on China vs. the American AI Labs (Ramp AI Index, 2026-07-08; empirical). Cited here for the instrument, not the substance: a payments platform's own administrative records, published as a monthly market index by that platform's lead economist. COI and sampling: Ramp measures its own corporate-card/bill-pay customers — a business-spend-active, VC-forward-skewed base, not a random sample of US businesses — and has a commercial interest in owning the authoritative AI-adoption dataset. The Google/Microsoft series used for the aperture argument comes from the raw file's recovered Datawrapper chart datasets (the letter names neither vendor); substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth, the open-weight demand reading at The Open-Weight Frontier Gap
  • The State of AI Impact in Engineering: Q2 2026 — Justin Reock, The State of AI Impact in Engineering: Q2 2026 (DX, Engineering Enablement newsletter, 2026-07-22). Evidence tier corrected empirical (as ingested) to vendor-claim at compile — see the Source Notes entry in Sources; in short, a developer-productivity vendor's lead-gen readout over a self-selected panel of its own 500+ customer organizations, mixing self-report with SDLC telemetry, with the methodology in a gated PDF the vault does not hold. Cited here for the instrument argument above (the opening control-group passage) and for the self-reported 34%-to-52% code share. Web article, no docling parse, so no table-collapse or table-shift discipline applies; the on-page charts are images and every figure quoted anywhere in this vault appears in the newsletter's own prose. Substance at Acceleration Whiplash and AI as Primary Author
  • AI-Augmented Human Resource Management? Insights from German companies — Kalff & Simbeck, arXiv 2607.13839 (2026-07-15 / v2 07-20; empirical, mixed methods, N=410 German HR managers). Cited here for the instrument, not the substance: §3 (survey design and the self-report limitation), §4.1 (the 183-of-410 informal-use finding and the boundary that company-specific AI "cannot be used informally"), §4.2 (the construct collapse — "did not recognize machine-learning systems… as 'actual AI'" — and the two-directional strategic use of the AI label by vendors and by firms facing co-determination), §5 (stated limitations). This document required a glyph-level repair at ingest — docling emitted every digit and every DOI/URL label as glyph names (/two_os, /D_SC), restored by a deterministic 1:1 substitution and re-verified against the PDF, so every number quoted from it traces through that repair. Table 1 is row-shifted and is cited nowhere; Table 2 verified clean. Full treatment and the complements reading at Organizational Complements to AI.
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 36
Related articles
  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…

  • Open Questions Backlog

    _456 actionable open questions across 205 pages · 107 predictions · 9 notes · 147 in progress · 69 watching (entities),…

  • Organizational Complements to AI

    The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…

  • Returns to Expertise in Agentic Coding

    Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…