H
Howardism
Plate IIAI Coding Practice中文HOWARDISM

Telemetry vs. Survey Measurement

Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (instrument agreement tracks shared data source, not construct), randomization as the only causal instrument, the survey arm's counter-case (shadow AI is invisible to telemetry), and the Ramp payment-rail aperture; first cross-family convergence: Anthropic passing OpenAI mid-2026; and the vendor-survey rule stated twice — third-party fielding and a stated n do not move the tier when the conclusion is the product and the method is gated.

Article metadata
Publication details
Published:June 17, 2026
Filed:Concept
Domain:AI Coding Practice
Reading:60 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Telemetry vs. Survey Measurement

Sources#

Summary#

Faros AI's methodological argument, and the basis for the most consequential conflict in its 2026 report: during rapid AI transformation, perception lags reality, so survey-based engineering research systematically misses the downstream damage that system telemetry catches in near-real-time. Faros draws its findings from engineering systems (task trackers, IDEs, static analysis, CI/CD, version control, incident management) rather than from how developers feel, and uses that distinction to directly contradict Google's DORA 2025 conclusions.

vendor-claim source — Faros's own platform is the telemetry instrument, so "telemetry beats surveys" is also a sales argument for that platform. The methodological point stands on its own merits, but the conclusion conveniently favors the vendor's product. See Acceleration Whiplash for the full evidence note.

Why perception lags reality#

The mechanism Faros proposes: at the individual level developers genuinely are more productive — task completion is up, code flows faster, the tools feel powerful — so surveys capture real, positive feeling. What surveys cannot capture is what happens downstream: "the review queues quietly backing up, the incidents accumulating in production, the bugs reaching customers." By the time those consequences show up in how people feel, "months have passed and the signal is already stale." Telemetry, drawn from the systems where work actually happens, does not lag. The claim: engineering leaders making consequential decisions about headcount, tooling, and process "need data as close to real time as possible… not how people feel about the work after the fact."

The DORA contradiction#

This is a flagged inter-source contradiction. DORA's 2025 State of AI-Assisted Software Development concluded that AI amplifies existing strengths and weaknesses, and that strong engineering foundations protect against AI's downsides. Faros's telemetry, it claims, "does not support that as a protective factor": high-performing organizations experience the same downstream deterioration as everyone else (see the maturity-independence finding in Acceleration Whiplash).

Weighing the conflict by method and incentive:

  • DORA 2025 — survey-based; large, long-running, vendor-neutral-ish (Google/DevOps Research). Strength: breadth and continuity. Weakness, per Faros: perception lag during fast transitions.
  • Faros 2026 — telemetry-based; within-company longitudinal comparison (low- vs high-adoption quarters), Spearman ρ at p<0.05. Strength: measures behavior, not feeling, near-real-time. Weakness: vendor-claim — Faros sells the platform, and "your mature practices won't save you, you need visibility + a context engine" is precisely the conclusion that grows its market.

Neither is a clean win. The honest read: Faros's measurement critique of surveys is sound (lagging perception is real), but its substantive claim that maturity offers zero protection should be held with the vendor incentive in view — it is the conclusion most favorable to selling the instrument. Worth tracking against future DORA editions and any non-vendor telemetry study.

The family effect: instruments agree with their data source, not with their construct#

The sharpest evidence that this page's dichotomy is a real fault line rather than a framing device comes from outside engineering metrics. Steele & Cruz (arXiv 2607.15506) put seven occupational AI-exposure instruments on the same O*NET occupations and correlate them pairwise. All seven claim to measure the same thing. What predicts whether two of them agree is which data source they were built from:

  • The two built from 2025 Anthropic Claude usage correlate at ρ = 0.89.
  • The two built from theoretical generative-AI capability (GPT-4 task ratings; 2,000 MTurk ability ratings) are the next-strongest pair — despite differing in level of analysis, rater, and question asked.
  • Across families, correspondence largely collapses; patent-mining, ML-rubric, and bottleneck-based instruments show "very little correspondence with each other or with later measures."

Two consequences for this page. First, the telemetry/survey split is not just a latency difference (behavior now vs. feeling later) — it is a partition of the answer space. Instruments in different families do not produce the same ranking, and in Steele & Cruz's case they do not even produce the same sign on the exposure-salary gradient. Second, the paper is the first entry in the vault where telemetry is run forward into a projection rather than reported as an observed count: query-volume ventiles rank the tasks, and a hand-set schedule of assumed automation ceilings supplies the levels. That is a hybrid — telemetry's ordering, assumption's magnitudes — and it inherits the credibility of the ranking without inheriting it for the numbers. See Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated for the full head-to-head.

The third instrument: randomization, and the one thing neither telemetry nor survey can do#

This page's dichotomy has a missing third term. Telemetry and surveys are both observational — one records behavior as it happened, the other records how it felt. Neither can say what would have happened otherwise, which is the question every claim on this page is ultimately about. A randomized experiment can, and Jabarian & Henkel (arXiv 2607.28222) is the vault's cleanest instance: 67,056 job applicants randomized between AI voice interviewers and human recruiters, pre-registered, with administrative employment outcomes.

What makes it useful here is that the paper runs all three instruments on the same event and they point in different directions:

InstrumentWhat it says about AI-conducted interviews
Administrative records (offers, starts, retention)AI-interviewed applicants receive 12% more offers, start more often, and are retained more (all p<0.001)
Applicant survey (post-interview CX survey)AI interviews are rated significantly less natural (p=0.014), pulling the composite perceived-quality index in the humans' favor (p=0.076)
Expert forecast (recruiter survey, pre-disclosure)61% of recruiters expected AI-led interviews to be of lower quality; 36% expected lower offer rates; 48% expected lower retention

The forecast is simply wrong — 131 experienced practitioners, surveyed before results were disclosed, predicted the sign backwards. The applicant survey is not wrong; it measures a real and different thing (naturalness) that does not track the outcome. And the administrative records are the only reason we know which is which.

Two consequences for this page. First, the felt-vs-system split it was built on is a special case of a bigger problem: an observational instrument is a claim about correlation regardless of how good its latency is, which is why CMU's 2.5M-PR telemetry study could support opposite conclusions from the same traces. Randomization is the rung above both, not a third flavor at the same level. Second, the honest cost of that rung: it took a firm partnership, an IRB, a pre-registration, and three months of a live hiring pipeline to answer one question about one task in one labor market. Telemetry and surveys are cheap and broad; the instrument that actually identifies the effect is expensive and narrow, which is exactly why the vault has hundreds of the former and one of the latter.

The survey arm, made concrete: what only asking can see, and what it costs#

Everything above argues against surveys. Kalff & Simbeck's N=410 study of German HR departments (arXiv 2607.13839, empirical) is the survey arm of this contrast in its plainest form — a self-reported adoption census of one business function, whose authors list "reliance on self-reported survey data for quantitative measures" among their own limitations. It is useful here for the opposite reason to the rest of the page: it contains the clearest instance the vault has of something no telemetry can see, and the clearest instance of the price of asking.

The thing only a survey can measure. Of the 410 valid respondents, 183 report using AI tools informally, on personal devices, regardless of employer policy or the availability of company-provided systems — including at companies whose official policy restricts or prohibits generative AI. The paper draws the boundary itself: "AI solutions requiring complex, company-specific data cannot be used informally, as they are accessible only through an official rollout." So corporate telemetry sees exactly the half of adoption that went through IT, and is blind by construction to a channel roughly as large. This is the CircleCI aperture problem in its limiting case — the instrument's coverage is whatever the vendor's product touches, and shadow adoption touches nothing the firm operates. On this question the ordering inverts: the survey is not the lagging instrument, it is the only instrument.

And the price of asking, in an unusually legible form. The same paper documents its own construct collapsing. Group discussions "often devolved into debates over the meaning of AI"; participants' perceptions were "strongly shaped by their exposure to generative AI applications popular in the media," so "some participants did not recognize machine-learning systems that underpin traditional HR analytics — such as those used to measure turnover risk or identify workforce trends — as 'actual AI.'" That is not ordinary measurement noise, because it is directional and it points at the paper's own headline: the conclusion is that predictive analytics plays only a minor role, and the tools respondents are least likely to count as AI are precisely the predictive-ML ones. A license- or spend-based instrument (Ramp's payment traces) has no such problem — it counts what was deployed, not what gets called AI.

Worse, the ambiguity is not merely fuzzy, it is contested by interested parties in both directions: vendors "leverage AI branding to generate interest and support business cases," while "the AI aspects of HR tools may be downplayed to avoid scrutiny from co-determination bodies." When the label is a strategic instrument for the people answering the question, no amount of question wording recovers the construct — a third failure mode, distinct from perception lag (this page's original argument) and from the family effect (instruments agreeing with their data source). Perception lag says respondents report an old truth; the family effect says instruments disagree about a stable construct; this says the construct itself is being moved by the respondents while you measure it.

The payment rail: the platform-records family gets a second member, and the first case where survey and behaviour agree#

The Connections entry below names Indeed's job postings as a fourth instrument family — a platform's by-product record of transactions it brokers, neither telemetry nor survey nor randomization. Ramp's monthly AI Index (empirical) is that family's second member, and putting the two side by side sharpens what the family is: whoever runs the rail can count what crosses it, for free, in near-real time, for exactly the population that transacts there. Indeed brokers vacancies and can therefore count labor demand; Ramp brokers corporate payments and can therefore count who firms pay for AI. Both are published by the operator's own research arm, so both fold the vendor incentive and the instrument into one party the way this page's Faros entry does.

The aperture bites in a checkable way here. Ramp's June-2026 vendor shares put OpenAI at 39.5% and Anthropic at 42.4% of businesses — and Google at 6.4% and Microsoft at 1.7%, with Google flat in the 4.6–6.4% band for the entire 42-month series while overall AI adoption went 7.5% → 55.0%. The most likely reading is not that Microsoft and Google lost the enterprise; it is that enterprise agreements, negotiated invoices and cloud committed-spend drawdowns do not cross a corporate-card rail the way a per-seat or per-API subscription does. That is the CircleCI aperture problem restated for payments: the instrument's universe is the purchases that fit its rail, and a vendor whose sales motion routes around that rail is structurally undercounted regardless of its actual share.

And the part this page did not previously have an example of: a survey and a behavioural instrument agreeing. The family effect above predicts that instruments track their data source rather than their construct, and most of this page is disagreements. Here two instruments from different families, on different populations, converge over the same six months on the same reordering: ICONIQ's Q2-2026 exec survey of ~305 AI-building software companies has Anthropic going 51% → 81% of respondents and passing OpenAI (77% → 71%), while Ramp's payment records have Anthropic passing OpenAI in May 2026 (42.4% vs 39.5% by June) with OpenAI down ~1.9pp from its November-2025 peak. Self-report and receipts, a builder cohort and a whole card base, same direction and roughly the same timing. Convergence across families is weak evidence taken alone and strong evidence taken here, precisely because this page's default expectation is that it does not happen — and it is the cleanest case in the vault of the instrument question being settled by agreement rather than adjudicated.

The convergence held through the next edition, and the rail showed a second aperture (September 2026). Ramp's September 9 letter puts August 2026 at Anthropic 43.8% against OpenAI 39.8% of businesses — the payment-rail lead widening from +2.9pp to +4.0pp while overall paid adoption reaches 56.1% — so the reordering the two instrument families agreed on is not a one-month artifact on the receipts side either. The same edition exposes a property of this instrument that belongs on this page, because it is an aperture in time rather than in coverage: Ramp's dollar-and-count series revise upward for months after first print (the letter restates July top-1% per-employee spend from ~$7.4K to ~$8.0K; June median per-employee spend reads $10.59 in the July edition and $10.94 in this one), while the binary vendor-share series barely moves. A rail counts transactions as they settle, so the freshest point of a spend series is the least complete one and any month-over-month change read off it is biased downward — the payments analogue of the revision problem administrative statistics carry and surveys do not. Instrument provenance at Ramp; the arithmetic at Firm AI-Spend Intensity and Headcount Growth.

"The control group is no longer viable" — a claim to split, not to accept (2026-08)#

DX's Q2 2026 AI-impact readout (vendor-claim, 500+ customer organizations) opens with the sharpest methodological assertion any source in this vault has made about its own instrument, and it is aimed squarely at this page's subject:

"When we first began tracking the impact of AI on engineering teams, our primary goal was to measure AI cohorts against historical baselines… With industry-wide AI adoption exceeding 90%, comparing AI users against a non-user control group is no longer a viable measurement strategy."

The claim is true of one control design and false of another, and the two are not usually distinguished. What saturates at >90% is organization- and developer-level adoption — whether a person or a firm uses AI at all. What has not saturated is the artifact-level mix: the share of individual changes, PRs or completions that an AI actually wrote.

  • The between-firms (or between-developers) adopter-vs-non-adopter contrast is genuinely dying, and DX is right about it. At >90% adoption the non-adopter arm is a residual of holdouts selected on whatever made them hold out, which is a worse confound every quarter. This is also the design DX's own panel and Faros's adoption-depth cross-section are built on, and the reason neither can separate "AI code is worse" from "orgs that adopt hardest differ in other ways."
  • The within-firm, within-codebase provenance contrast is alive and was running while DX declared it dead. Tran et al. (Google) compare AI-authored against human-authored changes inside one monorepo, same window, stratified on change size, over 3.52M submissions. Adoption in that population is effectively total — everyone has the tools — and the control cohort survives anyway, because AI's share of submitted code ran 28.99% to 68.62%, leaving roughly three in ten changes human-written at the end of the window. Universal adoption and a usable control cohort coexist without tension the moment the unit of comparison stops being the person.

So the correct statement is narrower and more useful than DX's: the control group moved down a level rather than disappearing. The population-level arm is gone; the per-change arm is not.

Why the claim is also structurally self-serving, in the way this page's Faros entry already documents. DX sells developer-productivity measurement, and its instrument is a survey-plus-SDLC-telemetry panel across customer organizations — an instrument that can only ever run the between-firms contrast. The design it declares non-viable is the one it cannot run; the remedy it proposes (longitudinal within-panel trends against historical baselines) is the product. That does not make the observation wrong, and the honest reading is that DX has correctly identified a real measurement crisis and misattributed its cause.

The cause is instrumentation, not adoption. A per-change control cohort requires authoring-time provenance — a record, made while the code is being written, of what wrote it. That is a decision taken before the artifact exists, by whoever owns the editor, and it is why Google could build both arms and nobody outside such a company has. The alternative available externally is a corpus defined by agent authorship (AIDev), which yields one arm by construction and no matched human comparison — which is exactly why Dipongkor et al. and Jhanglani et al. and Sakib et al. each measure a level and not a delta, and why DECODE can see only accepted completions. Four studies missing the same arm, for a reason that has nothing to do with adoption rates. Post-hoc AI-code detectors are the obvious substitute and Tran et al. reject them outright ("generalize poorly across models and settings").

The prescription that follows for reading this vault: when a source says it could not build a control group, ask whether it lacked non-adopters (increasingly unavoidable) or lacked provenance (a fixable instrumentation gap that most publishers of AI-impact numbers have not paid for).

The same instrument, regressed against itself (2026-09-22). DX's follow-up AI accelerates output, not innovation (Grace Fu, 2026-09-09, vendor-claim, 500+ customers; whether it is the identical Q2 panel is unstated) is a worked example of what a survey-plus-telemetry panel can and cannot find. Its headline outcome and its headline predictor are both respondent-reported — time saved (3.0 h/week in Q3 2025 to 6.1 h/week in Q2 2026) and an "AI output" composite whose AI-authored-code-share component the Q2 report carried as self-report (which components are telemetry is again unstated) — and that pair fits at 63% of variance explained. The outcome that is a different self-report, the innovation ratio (share of engineering effort on new capabilities versus maintenance), fits the full 15-metric model at 13%, with AI output at standardized β 0.16 and information-seeking friction at β −0.19 as the largest coefficients. The vault's reading, not DX's: the 63% is the fit to expect when predictor and outcome share a respondent and a frame — common-method variance — and the 13% is closer to what the constructs actually share. DX draws the substantive conclusion (AI accelerates output, not innovation) and recommends each customer re-run the regression on its own data; the methodological conclusion is that a panel with no provenance arm and no telemetry-side measure of the innovation ratio cannot separate "the hours do not convert into new features" from "the survey cannot see where the hours went." The two readings of the same coefficients are laid against each other on Acceleration Whiplash.

The vendor-commissioned survey, second instance: who holds the clipboard doesn't move the tier (2026-09-02)#

Prophet Security's State of AI in the SOC 2026 is the vault's second vendor-published survey of AI outcomes, after DX's State of AI Impact, and it is a useful test case because it is methodologically better than DX on every dimension the vault could check and still lands in the same tier.

What it has that DX did not:

  • A stated n — 250 IT and cybersecurity professionals — and a stated instrument length (31 questions).
  • A named third-party fielder: the research firm ViB screened and fielded it, and the gate page calls the work "vendor-neutral, 3rd party research… independently conducted by ViB."
  • Bases attached to the numbers in the write-up itself. The article distinguishes "of respondents", "of AI users", "among those who observed them", and "of teams that attempted a build" — the exact discipline whose absence made every DX figure uninterpretable.

What it still does not have:

  • No fielding dates, no question wording, no demographics, no per-question base counts. Methodology and demographics are named as contents of the gated report, behind a lead-gen form.
  • The conclusion is the product. Prophet sells an agentic AI SOC platform, and the survey's two most quotable findings — 46% of in-house builds scrapped or replaced, and an FAQ contrasting "an assistant embedded in one vendor's console" with "a purpose-built agentic platform" — are the buy-side pitch stated as data.
  • Every outcome figure is self-report about one's own team's performance, which is the instrument class this page's whole thesis says errs optimistically during a fast transition.

So the rule the two instances establish together: the tier turns on whether the conclusion is the product and whether the method is checkable — not on who administered the questionnaire. A neutral fielder improves sampling and screening; it does not make a seller's summary percentages reproducible, and "vendor-neutral" here is the vendor's own characterisation of research it commissioned.

The sharper point is what the survey arm is standing in for. In the coding domain this page can at least stage survey against telemetry — Faros against DORA, DX's self-reported code share against Google's byte-level provenance. In security operations there is no telemetry arm at all, and the absence is not technical. A SOC logs both of the quantities this survey asks people to recall: the time an investigation took, and the share of alerts nobody opened. Both are computable from a SIEM and a ticket system. The vault's entire population baseline for Autonomous Defense therefore rests on the one instrument family that cannot see the thing it is asked about — "72% of AI users report cutting investigation time by 25% or more" is a felt-time claim about an interval the respondent's own tooling timestamps. That is the cleanest available example of this page's aperture argument running in reverse: not an instrument whose coverage is limited, but a measurable quantity nobody has instrumented.

The self-report instrument at its largest: a decade series that changed its own question (McKinsey, August 2026)#

The state of AI in 2026 (McKinsey / QuantumBlack, 2026-08-25, empirical by tier, wholly self-reported) is the biggest survey instrument in this corpus — 1,719 participants in 97 nations, GDP-weighted, one respondent per organization, fielded May 4 - June 8 2026 — and the tenth annual edition of a series routinely quoted as the adoption baseline. Three properties make it a useful specimen rather than just another survey row.

1. It grades its own forecast, and fails. The 2025 edition reported 32% of respondents expecting AI-related headcount reductions over the coming year; the 2026 edition reports 14% saying reductions happened, with 66% reporting little or no change (Exhibit 16). A 2.3× overshoot, measured by the same instrument on the same panel one year apart. This is the rarest thing a survey can supply — a within-instrument calibration of its own forward-looking answers — and the direction matches the construct-collapse prior on this page only partly: here the expectation ran hot, not the retrospective report.

2. Its headline trend line is partly definitional. The 21% (2017) → 89% (2026) adoption series in Exhibit 1 carries a footnote conceding that the question changed four times: 2017 asked about AI "in a core part of the business or at scale"; 2018-19 about embedding at least one AI capability in processes or products; 2020+ about adopting AI in at least one function; 2025-26 about regular use in at least one function. That is the anchoring problem in the open question below, observed in the wild on the most-cited adoption series in business media — and observed without being quantified, since no year reports both definitions. Exhibit 6 carries a second, disclosed break: prior editions asked about cost impact at the use-case level and rolled up, 2026 asked at the business-function level directly, so the function-level cost figures are not comparable with earlier editions.

3. It measures attributions, not events. "AI has contributed positively to our EBIT" (37%) asks a respondent to make a firm-level causal judgment; "AI improved my productivity" (80%) is a felt measure of exactly the kind that diverges from system outcomes during a fast transition. The 6% high-performer construct is defined on the attributed outcome and then used to explain it. And the respondent frequently does not know: don't-know shares run 8-13% on the workforce items and reach 11% on the software-coding-agent phase question at large enterprises — on a panel where 654 of 1,490 role-identified respondents are C-level and 195 are individual contributors.

None of this demotes the tier: the fielding, sample and weighting are disclosed, which is more than most rows on this page. It places the instrument precisely — the widest aperture in the corpus, pointed at perceptions, with a ten-year trend line whose slope includes its own redefinitions.

A log carrying a change model, not an activity count (2026-09-22)#

Adoption Telemetry (Damon A. Young, PolyWise Partners; arXiv 2608.23617, 22 Aug 2026, 23pp) attacks this page's dichotomy from a third side. Tier it first: practitioner-opinion, single author, a consultancy that would sell the instrument, and no figure in it comes from a real deployment — the whole evaluation runs on synthetic populations the author generated. Read it for the argument, never for a number about the world.

The argument is that telemetry and survey are not two instruments for one construct but four traditions each instrumenting something adjacent to it. Observability instruments the system (traces, latency, cost, output quality — the human enters as a rater or a signal source). Enterprise usage dashboards instrument activity, and the example is exact: Microsoft defines an active Copilot user as one who "performed at least one intentional action in a supported application within the preceding 28-day window," so "an employee who samples an AI assistant twice a week and an employee whose work it has restructured are indistinguishable on a usage dashboard: both are active users." Product analytics owns the threshold craft, from the vendor's seat. Change management owns the construct — staged models of behaviour change, ADKAR's Awareness/Desire/Knowledge/Ability/Reinforcement — and "measures by asking": Prosci's 2026 AI Adoption Diagnostic is an expert-mediated assessment across five dimensions grounded in a survey of 1,100+ professionals. The proposed third thing is to express a change model's milestones as computable predicates over the event stream a deployment already emits, so adoption state is inferred from behaviour rather than reported — usage dashboards answer how much is it used, evals answer does it work, surveys answer how do people feel, and none answers where in the process of changing how work is done has this population stopped.

The definitional-sensitivity datum, secondhand and worth the aperture argument anyway. The motivating figures are ActivTrak Productivity Lab press material, carried as a citation and not checked against the underlying release: quarterly tracking of 120,620 workers across 1,009 organizations finds 82% of AI users sustain usage quarter over quarter while only ~2% reach a level of use consistently embedded in workflows, and a second release (443M work hours, 163,638 employees) finds 57% of AI users spend under 1% of working hours in AI tools. The shape is what counts here: one telemetry dataset yields 82% or 2% adoption depending only on where the depth threshold sits — a ~40× swing with the data held fixed. That is the construct-definition effect the anchoring question below asks a split-ballot survey to measure, appearing inside a behavioural instrument instead: the telemetry arm inherits the definitional problem rather than escaping it.

Two arguments this page did not previously carry. First, the aperture problem restated as governance rather than coverage: what a platform vendor structurally cannot supply is neutrality — its instrument "diagnose[s] its own product, in constructs it defines, revisable at its discretion, scoped to its own ecosystem," so a portfolio-neutral instrument that could support a substitution decision "is one no platform vendor has an incentive to build, and none has built." Microsoft's cross-organization benchmarking is the exception that shows the bound: by its own documentation those benchmarks cover "the percentage of active Copilot users" and at most 6 months of history — comparison exists for breadth alone. Second, the telemetry arm's own Goodhart exposure, which the vendor-survey entries above have no analogue for: "a population that knows its adoption is measured can perform adoption: invocations without reliance, long sessions without integration."

Where it lands on the survey arm. As a complement, on the AEI footing this page already describes: it "measures behavior, not experience," and its blind spots are survey-shaped (a zero-activity row cannot separate never heard of it from knows and declined; a long session cannot separate deep integration from struggle). §8.6 sharpens that into a research question: do perception instruments systematically lead telemetry as stall indicators, at what interval, and do they disagree in characteristic ways. Stage model and synthetic-only validation on Pilot-to-Production Gap; the classifier-free measurement design on Usage-Telemetry Classifier Validation.

Connections#

  • AI Adoption in Scientific Work — a tenth aperture: the survey pointed at permission rather than at behaviour, impact or feeling. Angelini & Lyrvall's latent class analysis of 3,785 PhD students asks how comfortable each respondent is with AI doing five named research tasks, and recovers four profiles from the pattern rather than from any single answer — an instrument that measures the shape of a person's acceptance instead of its level. Two properties earn it a slot here. It is the rare case where a behavioural measure sits inside the attitudinal instrument: AI-use frequency is a covariate, so the attitude-behaviour link is estimated rather than assumed, and heavy users turn out not to be across-the-board enthusiasts (daily/weekly use gives −1.911 on the refusing profile but a non-significant +0.434 on the permissive one). And it carries this page's construct problem in its own title: the paper calls the profiles normative boundaries while the items ask about comfort, and §5.4 concedes that comfort may equally reflect familiarity, access or beliefs about reliability. That is the split-ballot problem again — the finding is named for a construct the instrument does not isolate, and no version of the survey exists that asks legitimacy separately. The page it sits on also holds the matching telemetry, and the two disagree: 42.5% of science Gemini interactions are quantitative analysis, the task PhD students are least willing to delegate

  • Procedural Value in AI Decisions — an eighth aperture: the randomized stated-preference design. A conjoint is neither telemetry (nothing is observed) nor a survey in the attitudinal sense (the attributes are randomized, so the AMCEs are causal within the vignette), and it buys the one thing neither of the others can — orthogonal variation in procedure and performance, which no deployment generates because employers never randomize who decides against how often the system errs. Wang, Sturgis & de Kadt use it to price decision authority at +0.272 in choice probability against +0.285 for a 20-pp cut in wrongful rejections. It pays for that in external validity, and the same paper measures the exchange rate: intention to apply runs above belief in the system's legitimacy (3.24 vs 2.85, p<0.001, d_z=0.42), with 7.8% doubting the system yet intending to apply — a direct estimate of how far a behavioral proxy and a normative judgment separate when the respondent has no outside option

  • The Enablement–Regulation Axis — a ninth aperture: the elite-discourse text corpus. Not telemetry (no behavior is logged), not a survey (nobody is asked), not a conjoint (nothing is randomized) — a near-census of what one consequential population said on the public record, with a property none of the other apertures has: the speaker was institutionally obliged to be on record, so there is no non-response and no recall error. Chueri & Törnberg run it at 1,514,950 substantive speeches across 33 national parliaments (2023–Apr 2026) and read frame composition off it. What it buys is completeness and latency on the supply side of policy; what it pays is construct — a speech measures positioning, not preference, not behavior, and explicitly not enactment ("parliamentary speech does not directly measure enacted policy"). This page's definitional-sensitivity mechanism appears here as the retrieval dictionary rather than a threshold: 1.5M speeches are cut to 19,411 candidates by multilingual keyword lists whose recall is never estimated, so what counts as an "AI and work" speech is a boundary the pipeline draws once and never varies — the split-ballot design this page keeps asking for, unrun again

  • The Enterprise AI Adoption Gradient — a seventh aperture: the vendor's administrative ledger. OpenAI's enterprise study opens by arguing this page's case — survey self-report "is typically less detailed and may suffer from imperfect recall or reporting biases" — and then demonstrates the trade in both directions. Account records resolve depth inside buyers (messages, WAU and output tokens per organization-week, linkable to Compustat) at a precision no survey reaches, and they are structurally blind to non-adoption, rival vendors, shadow use and any outcome at all — the same blindness this page records for the Ramp payment rail, one product narrower

  • Closed-Loop AI Review — the definition problem on mined artifacts rather than on survey items. This page's mechanism is that a measured rate moves with the threshold used to define the thing; there, the count of "agent-authored PRs" moves by 1,733,535 on one choice — whether a branch-name prefix like codex/ counts as evidence of agent authorship. Selvanayagam & Ghaleb decide it does not (96.9% of their entire 38.0% author-side quarantine), which is the single largest methodological decision in the study and the denominator of its headline 8.8% AI-review rate. The asymmetry is the instructive part: quarantine is 0.01% on the reviewer side, where vendors post under controlled bot logins, so a mined rate can be near-complete in its numerator and badly incomplete in its denominator at the same time

  • Agent Documentation Behavior — trace telemetry taken to its honest limit, and a rare case of a study deleting a dimension rather than estimating it. Gao & Chen's initial coding scheme included purpose — why the agent opened the file — and they removed it: purpose is not recoverable from tool-call logs, a file read is compatible with many intents, and assigning one "would be unfalsifiable." They report trigger, interaction type and outcome instead, and state the observation scope explicitly (repository-local file operations only; no browser reads, no model weights, no in-source docstrings; runtime-injected context files invisible until re-read, so every absolute rate is a lower bound). That is the clearest statement in the corpus of the boundary this page argues about — logs answer what happened, and the intent question they cannot answer is precisely the one surveys are reached for

  • Pilot-to-Production Gap — the diagnosis this page's instruments cannot locate. That page carries the same source's five-stage stall taxonomy and its synthetic-only boundary; the split is that the stage model is a claim about enterprise deployment failure and lives there, while "a production log can carry a change-management construct at all" is a claim about instruments and lives here. The pair shows the literature's gap from both ends: the pilot-failure documents diagnose organizational causes without measuring them, and this source proposes the measurement with no real deployment behind it

  • Autonomous Defense — the domain where this page's dichotomy collapses to one arm. Security operations has the logs to measure investigation time and uninvestigated-alert share directly, and the corpus's only population reading of either is a vendor-commissioned self-report survey (n=250, fielded by ViB) — so the SOC's baseline numbers carry the survey arm's optimism skew with no telemetry counterpart to check them against

  • Community Smells Under AI Adoption — the instrument tension appearing inside one instrument: the same survey's free-text layer reports AI displacing teammates while its structural models estimate peer interaction rising. The authors' three-part reconciliation — individual change vs cross-respondent association, concentrated-and-conditional effects, and a non-significant direct path indicating mediation rather than absence — is a reusable frame for reading self-report here

  • The Open-Weight Frontier Gap — the limit of a payment-rail instrument on the question it is most often used for: Ramp's 5.8%-of-AI-spenders model-serving figure is the vault's main quantitative handle on open/Chinese-model adoption, and it can only see firms that pay a serving vendor — downloading open weights and running them on your own GPUs generates no transaction and is invisible by construction, the same blind spot shadow AI creates for corporate telemetry above

  • Organizational Complements to AI — the substantive home of this source, and the reason its instrument matters: adoption below the organizational level (183 of 410 using AI outside any rollout) is a complements story that firm-level telemetry cannot reach, so the complements literature's usage-share and spend-trace instruments are structurally blind to the fastest-diffusing half

  • Controlled Variance: AI's Edge as Reduced Dispersion — the third instrument above: the vault's only randomized causal estimate of AI substituting for a human in an expert conversational task, and a single design in which administrative records, an applicant survey, and an expert forecast disagree about the same intervention

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — where the seven-instrument comparison lives; the family effect above is the measurement-methodology reading of it, and its telemetry-built seventh instrument is the hybrid case this page's dichotomy did not previously have a slot for

  • Usage-Telemetry Classifier Validation — the price of the telemetry side of this argument: system logs beat self-report on latency and scale, but the economic claims drawn from them run through an LLM classifier whose exact accuracy (22.6% at the O*NET task level) was never published until Google ATLAS did it

  • AI Native Product Cadence — the limit of the telemetry side: OpenAI's Akshay Nathan reports the classic engineering proxies (commits, LOC, PRs, tokens) decoupling from goal attainment under agents, so better instrumentation of the same signals measures motion rather than progress — telemetry beats survey on latency, but only if the logged quantity still tracks the outcome

  • Acceleration Whiplash — the maturity-independence finding rests on this telemetry-over-survey methodology

  • Production-Sourced Evaluation — the same "measure from the real system, not a proxy" instinct applied to model evals; telemetry-vs-survey is its engineering-metrics cousin

  • Evals as Product Spec — Cat Wu's evals encode the spec; telemetry encodes what actually shipped — both prefer ground-truth signal over self-report

  • Verification as the New Bottleneck — Fiona Fung's warning to break PR-cycle-time into funnel chunks rather than read the aggregate is the same "instrument carefully or the signal misleads" discipline. That page also carries the instrument-aperture case: CircleCI's Q2 2026 Pulse (vendor-claim, 20M+ workflows) is behavior-not-feeling telemetry like Faros's, but from a single platform on a single signal — it sees CI workflow runs and nothing else. That aperture fixes what it can claim: it counts validation cycles (its Merge Efficiency Ratio) with a precision no survey could reach, and cannot observe a production incident, a customer-facing bug, or a review queue at all. So its cheerful main-branch success rate (70.8%→76.7%) and Faros's grim incidents-per-PR (+242.7%) are barely in conflict — each vendor reports the half of the SDLC its own product instruments. The corollary for this page's thesis: telemetry beats surveys on latency, but its coverage is whatever the vendor's product happens to touch, and a vendor's headline metric is reliably one its instrument can see and its product can improve

  • Compounding Data Moat — owning the telemetry stream is itself a moat; the report is a demonstration of what the data asset enables

  • Agentic Coding Work-Composition Shift — the empirical cousin: Anthropic's Clio-based 400K-session telemetry reads behavior-not-feeling the same way, but is a research artifact (validated classifiers, controls) rather than a vendor-claim lead-gen report — and its session-layer optimism vs Faros's org-layer pessimism is the felt-vs-system split this page names

  • Conversation-to-Delegation Shift — the third major usage-telemetry study (OpenAI/Codex, empirical), and it extends this page's argument: as usage becomes delegation, even interaction-count metrics (active users, chats) go stale — track complexity, runtime, concurrency, reuse, output instead

  • Anthropic Economic Index — the program that resolves this page's dichotomy: it links usage telemetry to survey responses per person (the Cadences report, ~9,700 linked respondents), treating telemetry and survey as complements rather than rivals

  • AI Usage Cadences — the AEI's continuous hourly telemetry is this page's "measure the real system finely" principle pushed to time resolution

  • Review as the Control Point — the non-vendor telemetry this page's open question asked for, and a deepening of its methodology argument: CMU re-scraped 2.5M+ GitHub PRs (behavior, not feeling) yet found the same traces support opposite conclusions under defensible analysis choices — so telemetry beats surveys on latency, but surface telemetry can't adjudicate why without a causal model (Pearl: "data are profoundly dumb"). Telemetry's edge is real and bounded

  • Market-Priced AI Exposure (the AI Premium) — the purest realized telemetry (every observation is a paid request, behavior not feeling), pushed to cross-provider breadth (400+ models, ~2% of global tokens) and then turned into a market-priced signal via equity-price comovement; the finance-side entry in this thread — but a reminder that telemetry's own collection mechanism biases it (OpenRouter is developer-skewed), so realized ≠ representative

  • AI Investment Story, Not Efficiency Story — this page's dichotomy carried into startup economics: on revenue-per-employee, Emergence's measured cap-table receipts (AI companies below non-AI peers) conflict with AWS's self-reported founder survey (AI-natives above the baseline) — a survey-vs-financial-telemetry instrument split, same felt-vs-system shape as the Faros-vs-DORA case, resolved the same way (weight the receipts, flag rather than average)

  • Firm AI-Spend Intensity and Headcount Growth — also the home of a fourth instrument family this page had no slot for: platform administrative records, in Indeed Hiring Lab's job-postings series. Not telemetry (nobody's usage is logged), not survey (nobody is asked), not randomization (nothing is assigned) — it is the by-product record of transactions a platform brokers, which makes it behavior-not-feeling and near-real-time on the demand side, where usage telemetry sees only the supply side. It inherits the aperture problem in a sharper form than CircleCI's: the instrument is one job board, so its universe is whoever posts there, and "US job postings fell 7%" is a statement about Indeed's marketplace before it is a statement about the economy. It also collapses the vendor-incentive and the instrument into one party in the way this page's Faros entry does — Indeed's research arm publishing a rebound story about Indeed's own board — while the substantive limit is different from Faros's: the numbers are administrative and hard to dispute, and it is the causal reading ("Claude Code launched, then postings rebounded") that the design cannot carry. Beyond that, the same behavior-not-feeling instinct applied to AI adoption and its labor effect: Ramp reads adoption off actual AI-vendor payment traces (not "do you use AI?" surveys, whose answers for the same period span 18%→78%) and joins them to Revelio workforce records — a revealed-adoption instrument the paper positions explicitly against "messy surveys and exposure measurements"

  • AI Product Economics Maturation — the survey axis pushed to its softest form: ICONIQ's exec survey layers forward projections on top of self-report (2026P/2027P margin and RPE), so its rosy trajectory is felt-about-the-future, two removes from telemetry. Where its projected margin expansion (→59% by 2027) meets Emergence's measured growth-margin compression, this page's prescription applies unchanged — weight the receipts, flag rather than average, and mark the projection prediction-grade

  • Outsource Your Thinking, Not Your Understanding — the reverse case to this page's thesis: comprehension debt is damage the telemetry misses, because the shipped artifact is fine (Shopify reports AI-assisted reversion rates at pre-AI baselines) and the human is what changed

  • Standardize the Infrastructure, Not the Tools — the instrument that makes org-wide AI telemetry possible at all (one gateway, one meter), and the open question of what per-team usage analytics get used for once they exist

  • Efficiency Debt of AI-Generated Code — first-party telemetry with a control cohort, which is the shape this page's whole Faros argument has been missing. Google instruments its own monorepo at byte-level authoring provenance and compares AI-authored against human-authored code inside the same window and size stratum — so unlike Faros's low-adoption-vs-high-adoption cross-section it can separate "AI code is worse" from "orgs that adopt hardest differ in other ways." Two lessons for this page. The aperture rule holds and cuts deeper than usual: the instrument sees everything inside one company's pipeline and nothing outside it, which is total coverage of a population of one. And the vendor incentive runs the opposite direction from every other entry here — the org measuring the code is the org that built the tools that wrote it, and the result is on balance reassuring (revert rate below parity) with a discussion attributing the residual weaknesses to "historical default system configurations." Two telemetry sources, two directional COIs, opposite signs, and only one of them has a control group

  • AI and Market Power — the fifth instrument family, and the only one with a real claim to representativeness: compulsory national statistical surveys run by statistical offices to common Eurostat guidelines (France's TIC, Portugal's IUTICE), linked to administrative balance sheets. It is a survey, so it inherits the self-report limits this page catalogues — but it is a sampled, weighted, mandatory one rather than a self-selected panel, and it anchors the low end of the vault's adoption spread: 20.2% of OECD enterprises using AI in 2025, next to Census BTOS's 18%, Ramp's ~55% of eligible businesses on the payment rail, and executive surveys at 69–78%. The aperture argument gets its cleanest illustration here — five instruments, one phenomenon, a 4× spread, and the most representative one reads lowest

  • Psychological Costs of AI Adoption — a seventh instrument family, and the only one whose subject matter no other family has a column for: semi-structured interviews plus member checking (N = 21, one firm, plus 12 participants reviewing a 15-page findings report against their own experience). What it can support: appraisal — why practitioners do what the telemetry records them doing. It is the only instrument in the vault that can distinguish "read the diff line by line because the diff was risky" from "read the diff line by line because I am accountable for it," which is a causal claim about review load that no PR-level metric can reach, and it is the only one that observes absorbed strain at all ("I don't. I just keep going"). What it cannot support: magnitude, trend, or population inference. Its 10/10-of-software-engineers figures count participants whose transcript contained at least one matching code, which is a coding-procedure artifact reported honestly as "descriptive... exploratory rather than conclusive," and the sample is 21 volunteers reached through the employer's own AI program director, in one firm, at one time. Member checking is a real reliability gain over an unvalidated survey and is not external validity. Read alongside the free-text layer of Community Smells Under AI Adoption, it also repeats this page's family effect one level down: the two self-report instruments agree with each other about what practitioners feel while disagreeing with each other about what it implies for the team

  • Post-Acceptance Edit Behavior — a sixth instrument family, and the only one that observes an artifact which is then destroyed: pre-commit editor telemetry. DECODE saves the working file every time a developer pauses for one second, so it sees the AI code that gets deleted 23 minutes after acceptance — code that exists in no repository, no PR, no survey response and no monorepo history, and is therefore invisible to every other instrument on this page. That is the aperture argument's mirror image: the layers below the commit are not a smaller view of the same population, they contain events the higher layers structurally cannot record. The weakness is the familiar one in a sharp form — the aperture is one opt-in VS Code extension whose users installed a model-comparison tool, self-selected in a way a mandatory statistical survey or a company's full monorepo history is not

Derived#

Open Questions#

  • Is there a non-vendor telemetry dataset large enough to adjudicate the maturity-protection question independently of Faros's commercial framing? Partially answered: CMU's arXiv 2607.07980 supplies exactly this — a non-vendor, 2.5M+-PR GitHub telemetry study — and it (a) finds the agent no-review rate converging toward the human baseline rather than a widening gap, and (b) argues the effect's sign is set by team practice, closer to DORA. The catch: its own headline is that the telemetry is direction-unstable, so it counterweights Faros without cleanly settling the maturity question — the honest verdict is "surface telemetry alone, vendor or not, can't adjudicate this." The worked example of that verdict is The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence: the two telemetry studies' opposite under-review headlines dissolve once metric, population, time axis, and authorship unit are aligned — the datasets never disagreed, only the framings did.

  • Does anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction? Kalff & Simbeck document respondents excluding predictive ML from the category while counting ChatGPT in it, which would bias reported adoption down for exactly the tool class their headline finding is about. Falsifiable two ways: a split-ballot survey where one arm names the underlying technique ("software that scores turnover risk") instead of the label, or validating firm-level self-reported adoption against vendor-license/spend records for the same firms — the Ramp instrument applied as a criterion rather than a substitute. Related evidence (2026-08-12), not an answer: DX reports a self-reported AI-generated code share of 34% to 52% across Q1-Q2 2026 in 500+ organizations, while Google's authoring-time provenance over the overlapping window reads 28.99% to 68.62%. Different populations (a cross-industry customer panel vs one C++ monorepo), different constructs (code share vs adoption), no matched firms, and DX's question wording sits in a gated PDF the vault does not hold — so this settles nothing. It is nonetheless the corpus's first side-by-side of a self-reported code share against a provenance-measured one, and the direction is the one Kalff & Simbeck's construct-collapse mechanism predicts: self-report reads below the instrument that counts bytes, not above it. The criterion validation this bullet asks for is that same comparison run on the same firms. A second instance of the mechanism, documented but unquantified (2026-09-22): The state of AI in 2026: On the road to ROI publishes the most-cited AI adoption series in business media — 21% in 2017 to 89% in 2026 — with a footnote conceding the question was redefined four times over that span (2017 "a core part of the business or at scale"; 2018-19 embedding at least one AI capability in processes or products; 2020+ adopted AI in at least one function; 2025-26 regular use in at least one function). Each revision loosens the threshold, so an unknown share of the series' slope is the definition moving rather than the behaviour. The same edition discloses a second break: prior years asked about AI's cost impact at the use-case level and rolled up, 2026 asked at the business-function level directly. This is the bullet's mechanism caught in the wild at the largest scale available — and it is still not the answer, because no year is fielded under two definitions, so the effect size the split-ballot design would measure is exactly what remains unmeasured. What it adds is that the question is not hypothetical: a decade-long trend line in wide circulation already contains it. The same mechanism on the telemetry side, which this bullet assumed was immune (2026-09-22): Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals reports one behavioural dataset (ActivTrak, 120,620 workers) reading 82% adoption under a sustained-usage definition and ~2% under a workflow-embeddedness one — same logs, ~40× spread, definition alone. Secondhand and practitioner-opinion, so not the criterion validation asked for; it does mean the implied remedy (validate self-report against telemetry) must pin the telemetry threshold first, or inherit two definitional degrees of freedom instead of one.

Resolved Questions#

  • Surveys and telemetry measure different things (felt productivity vs. system outcomes); is the "contradiction" partly a category error — both true at their own layer — rather than one being wrong? Answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — yes for the measurement halves: surveys capture felt productivity (genuinely real at the individual layer — task completion did rise) while telemetry captures system outcomes that haven't propagated into feeling yet; during a fast transition the layers legitimately diverge, and the AEI's linked telemetry+survey design already treats them as complements. But not entirely: the maturity-protection claim is one proposition about one layer (system outcomes) and remains substantively contested — Faros's "no protection" aligns with its vendor incentive, CMU's moderator theory argues the sign is team-set (closer to DORA) but measures no maturity effect. The category-error dissolution cleans the framing; the one real disagreement stays open (tracked in the non-vendor-dataset question above).

Sources#

  • From Agent Behaviour to Agent-Friendly Documentation — Gao & Chen (Peking University), arXiv 2608.20195, 2026-08-20, empirical. Cited here only for §3.3's decision not to code interaction purpose, §3.5's Observation-scope paragraph, §7.1's lower-bound construct-validity argument and §3.6's cluster-bootstrap treatment of nested observations. A methodological citation, not a measurement one — the paper reports no survey arm and no productivity quantity. Full treatment on Agent Documentation Behavior

  • AI Engineering Report 2026: The Acceleration Whiplash — "A direct counterpoint to DORA's 2025 findings"; Research Methodology; Report's Purpose

  • DORA, 2025 State of AI-Assisted Software Development (cited by Faros): https://dora.dev/research/2025/dora-report/

  • 5 takeaways from the State of Software Delivery Q2 Pulse report — Jacob Schmitt, CircleCI blog (2026-07-08), vendor-claim: the single-platform telemetry aperture and the Software Delivery Data Explorer benchmark framing

  • Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, arXiv 2607.28222 (2026-07-30), empirical, pre-registered field experiment. §3.1 (administrative outcomes), §3 intro (the recruiter forecast: 36% lower offers / 48% lower retention / 61% lower interview quality), §5.2 (the applicant survey's naturalness result and the composite quality index). Full treatment and parse warnings at Controlled Variance: AI's Edge as Reduced Dispersion.

  • Helping People Choose Careers in the Age of AI — Steele & Cruz, arXiv 2607.15506 (2026-07-16), empirical. §4 (the query-based construction) and §4.4 + Fig. 6 (correspondence among the seven instruments). Table 1 and Table 2 verified clean at ingest; Table 5 and appendix Table A1 are damaged in the parse and are not cited here. The paper contains a source-internal contradiction in its own summary statistics — see the Sources note on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated.

  • AI and Job Postings: From Destruction to Creation? — Guillermo Gallacher, AI and Job Postings: From Destruction to Creation? (Indeed Hiring Lab, 2026-07-08; empirical). Cited here for the instrument rather than the substance: a platform's own administrative transaction records (job postings), analyzed and published by that platform, with an uncontrolled before/after design around a product launch. Substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth; chart-only quantities in the source (per-country shares, scatter-plot correlation statistics) are not quoted anywhere in this vault.

  • Ramp's latest data on China vs. the American AI Labs — Ara Kharazian, Ramp's latest data on China vs. the American AI Labs (Ramp AI Index, 2026-07-08; empirical). Cited here for the instrument, not the substance: a payments platform's own administrative records, published as a monthly market index by that platform's lead economist. COI and sampling: Ramp measures its own corporate-card/bill-pay customers — a business-spend-active, VC-forward-skewed base, not a random sample of US businesses — and has a commercial interest in owning the authoritative AI-adoption dataset. The Google/Microsoft series used for the aperture argument comes from the raw file's recovered Datawrapper chart datasets (the letter names neither vendor); substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth, the open-weight demand reading at The Open-Weight Frontier Gap

  • AI accelerates output, not innovation — Grace Fu, AI accelerates output, not innovation (DX newsletter, 2026-09-09; vendor-claim). Cited for the instrument, not the finding: self-reported time saved (3.0 → 6.1 h/week) against an AI-output composite fits at R² 0.63, the survey-derived innovation ratio against 15 workflow metrics at R² 0.13; β 0.16 (AI output) and −0.19 (information-seeking), both p < 0.01. Paraphrased-digest raw; numbers preserved, prose not quoted

  • The State of AI Impact in Engineering: Q2 2026 — Justin Reock, The State of AI Impact in Engineering: Q2 2026 (DX, Engineering Enablement newsletter, 2026-07-22). Evidence tier corrected empirical (as ingested) to vendor-claim at compile — see the Source Notes entry in Sources; in short, a developer-productivity vendor's lead-gen readout over a self-selected panel of its own 500+ customer organizations, mixing self-report with SDLC telemetry, with the methodology in a gated PDF the vault does not hold. Cited here for the instrument argument above (the opening control-group passage) and for the self-reported 34%-to-52% code share. Web article, no docling parse, so no table-collapse or table-shift discipline applies; the on-page charts are images and every figure quoted anywhere in this vault appears in the newsletter's own prose. Substance at Acceleration Whiplash and AI as Primary Author

  • AI-Augmented Human Resource Management? Insights from German companies — Kalff & Simbeck, arXiv 2607.13839 (2026-07-15 / v2 07-20; empirical, mixed methods, N=410 German HR managers). Cited here for the instrument, not the substance: §3 (survey design and the self-report limitation), §4.1 (the 183-of-410 informal-use finding and the boundary that company-specific AI "cannot be used informally"), §4.2 (the construct collapse — "did not recognize machine-learning systems… as 'actual AI'" — and the two-directional strategic use of the AI label by vendors and by firms facing co-determination), §5 (stated limitations). This document required a glyph-level repair at ingest — docling emitted every digit and every DOI/URL label as glyph names (/two_os, /D_SC), restored by a deterministic 1:1 substitution and re-verified against the PDF, so every number quoted from it traces through that repair. Table 1 is row-shifted and is cited nowhere; Table 2 verified clean. Full treatment and the complements reading at Organizational Complements to AI.

  • State of AI in the SOC 2026: 8 Key Takeaways — Ajmal Kohgadai (Director of Product Marketing, Prophet Security), State of AI in the SOC 2026: 8 Key Takeaways, 2026-08-03 (page JSON-LD says 08-04). Evidence tier vendor-claim, confirmed as ingested and not upgraded despite a stated n and third-party fielding — reasoning in Sources, following the DX precedent above. Cited here for the instrument comparison: n=250, 31 questions, screened and fielded by ViB, bases stated per figure in the write-up, methodology and demographics behind a lead-gen form, and a seller whose product is the survey's conclusion. Web article, no docling parse, every figure in prose rather than in a chart image. Substance at Autonomous Defense

  • The state of AI in 2026: On the road to ROI — Dan Tinkoff, Lieven Van der Veken & Michael Chui with Tara Balakrishnan, The state of AI in 2026: On the road to ROI (McKinsey / QuantumBlack, 2026-08-25, empirical, self-reported; online survey, 1,719 participants in 97 nations, fielded May 4 - June 8 2026, GDP-weighted). Cited here for the instrument, not the substance: §"About the research" (panel, fielding window, GDP weighting), Exhibit 1's definition-change footnote, Exhibit 6's use-case-to-function question change, Exhibit 16's expectation/realization pair, Exhibit 5's role bases (654 C-level / 424 executive or senior manager / 217 midlevel manager / 195 individual contributor), and the don't-know shares in Exhibits 3 and 16. Web article; all 18 exhibits are rasterized SVG charts transcribed at ingest, and Exhibit 11 carries no printed labels (gridline-read, approximate — not cited here). COI: McKinsey sells AI transformation consulting.

  • September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 — Ara Kharazian, September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 (Ramp, 2026-09-09; empirical). Cited here for the instrument again rather than the substance: the August vendor shares that extend the cross-family convergence, and the edition's own disclosure of an upward revision to a previously published figure, which is the evidence for the revise-after-print property described above. Same COI and sampling caveats as the July row; instrument page at Ramp

  • Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals — Damon A. Young (PolyWise Partners, sole author), Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals, arXiv 2608.23617 (2026-08-22), 23pp, practitioner-opinion. Cited here for the instrument argument only — §2.1–2.5 (the four traditions and the gap), §3 (the definition and its four properties), §6.5 (Goodhart; behaviour-not-experience), §7 (neutrality), §8.2 (the Microsoft benchmarking bound), §8.6 (perception as leading indicator). COI: a consultancy paper proposing an instrument its author's firm would deploy, closing with a design-partner invitation. Every quantity in it is synthetic or secondhand — the ActivTrak figures arrive via a PR Newswire release in the references, unverified against the underlying report; Table 1 was reconciled cell-for-cell against pdftotext -layout and is exact. Stage model and validation boundary at Pilot-to-Production Gap.

  • AI-to-AI Code Reviews of GitHub Pull Requests — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21, 15pp, ESEM 2026 Emerging Results track, empirical. Cited here only for §3.2–3.3 — the S1/S2 signature tiers, the decision to exclude branch-name evidence, and the resulting quarantine asymmetry (1,733,535 of 4,563,819 candidate author-side PRs, 38.0%, against 0.01% on each review stream) — as the artifact-mining instance of this page's definition-threshold mechanism. All five tables reconciled against pdftotext -layout at compile. Full treatment on Closed-Loop AI Review

  • Normative boundaries of AI in scientific work: Evidence from PhD researchers — Angelini & Lyrvall (independent / Inria), Normative boundaries of AI in scientific work: Evidence from PhD researchers, arXiv 2608.25678, 2026-08-26, 15pp, empirical. Cited here for the instrument only: §3 (the five comfort items, the self-selected 3,785-respondent Nature Graduate Survey 2025 frame), §4.1 (AI-use frequency as a covariate on attitude class) and §5.4 (comfort is not legitimacy). Substance, the four profiles and the telemetry contradiction on AI Adoption in Scientific Work

§ end
Cited by 49
Related articles
  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…

  • Organizational Complements to AI

    The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…

  • Returns to Expertise in Agentic Coding

    Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…