Sources#
- 5 takeaways from the State of Software Delivery Q2 Pulse report
- AI accelerates output, not innovation
- AI Engineering Report 2026: The Acceleration Whiplash
- Characterizing the Quality Profile of AI-Generated C++ in Production
- Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild
- The Speed Trap: 8 takeaways from our latest AI engineering research
- The State of AI Impact in Engineering: Q2 2026
- When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering
Summary#
The central finding of Faros AI's AI Engineering Report 2026 (telemetry from 22,000 developers across 4,000 teams, analysis as of March 2026): AI has flooded a system built around human-paced development and human-quality code with output it was never designed to absorb. Throughput rises sharply while quality degrades downstream, and — the report's load-bearing claim — the gap between the two widens as adoption deepens, rather than stabilizing. Faros names this the Acceleration Whiplash: "the acceleration is real, but it is deceptive — it masks the strain building at every stage downstream."
Evidence note. This is a
vendor-claimsource — Faros sells an engineering-intelligence platform and the report's prescription (a "context engine," recommendation #10) maps to its product category. The underlying data is genuine telemetry (Spearman ρ, p<0.05, within-company over time), so the measurements are empirical, but selection and framing serve a commercial narrative. Claims below are attributed to Faros, not stated as settled fact. It also directly contradicts DORA's 2025 survey findings — see that page.
The two halves, quantified#
Throughput is up (low→high AI adoption, within-company):
- +33.7% task throughput per developer; +66.2% epics completed per developer
- +210% code-specific tasks completed per team (≈6× the general-task rate)
- +16.2% PR merge rate (but down from +98% in Faros's 2025 report — Faros reads the gap as a review bottleneck throttling merges)
- −11.7% deployments per week (10% of dataset); +861% code churn (lines-deleted-to-added ratio)
Quality is down, across every downstream stage:
- Cognitive load (see AI Brain Fry): daily PR contexts/dev +67.4%, work restarts +13.8%, stalled in-progress tasks (no activity 7+ days) +26%
- Complexity / wider change blast radius: avg PR size +51.3%, files edited per PR +59.7%, files touched per dev/month +149.9%
- Pre-merge quality: review comments +25%, PRs merged with no review +31.3% ("the most urgent finding")
- Flow: time-in-progress +225.2%, median time-in-PR-review +441.5%, lead time commit→prod +480.4% (10% of dataset)
- Production: incidents per PR +242.7% (probability of an incident per merge more than tripled), monthly incidents +57.9%, bugs per developer +54% (up from +9% in 2025), reopened tickets +12.6%
The maturity-independence finding#
Faros's most striking claim: the whiplash appears regardless of baseline engineering maturity. Organizations with strong pre-AI performance — mature DevOps, high DORA scores, disciplined delivery — see the same downstream deterioration as everyone else. "Even the strongest foundations are buckling under the landslide of AI-generated output." This is the explicit empirical wedge against DORA 2025, which concluded strong foundations protect against AI's downsides.
The thesis: it's an authoring problem, not a review problem#
The report's punchline reframes the fix. The natural instinct — more reviewers, stricter gates, longer QA — "treats the symptom." Faros argues the problem must be addressed at the source, during code generation: "the goal should be fewer mistakes arriving at review, not more humans deployed to catch them." AI-generated code is superficially convincing (idiomatic, well-named, stylistically consistent) while its structural failures sit beneath the surface, so it imposes a disproportionate tax on senior engineers — the only people equipped to catch intent-level errors, now consumed unraveling plausible-looking code that "was never ready."
This is a productive refinement of Verification as the New Bottleneck: Faros agrees verification is the binding constraint, but argues you relieve it by raising authoring quality (richer context at generation time), not by scaling the verification layer. The mechanism it prescribes — give agents codebase standards, architectural intent, security constraints, and a "context engine" built from how the codebase evolved, not its current state — is the industrial-scale version of the persistent-context discipline.
Why "whiplash" and not "paradox"#
Faros's July 2025 report named the AI Productivity Paradox (investment up, delivery gains not materializing). The 2026 report claims the paradox "sharpened into a crisis": adoption accelerated, the absorb-gap widened, and the throughput gains are real but front-loaded — they "mask the strain" that surfaces downstream weeks-to-months later. The whiplash is the temporal structure: fast visible acceleration, delayed invisible cost. Note the comparison is directional only — the 2025 and 2026 datasets are independent cross-sections, not a longitudinal panel.
A near-term boundary condition#
Faros stresses these numbers reflect AI as a primary authoring tool with humans still in the loop — agentic authoring is <1% of PRs in this dataset (see AI as Primary Author). "Remove that human from the loop entirely, and every metric here faces pressure an order of magnitude greater. The industry is not ready for that transition." The whiplash, on Faros's telling, is the mild version.
The split result: a human-controlled production measurement disagrees, but not everywhere (2026-08)#
Tran et al. (Google, arXiv 2608.06640) is the first source in the corpus that can be held against this page on its own terms: empirical rather than vendor-claim, production scale (3.52M submitted changes, April 2025-April 2026), and — the thing Faros lacks — a human-written control cohort, with outcomes stratified on change size, month, and organizational slice.
What it corroborates. Review friction is real and points the way this page says. AI-generated changes draw 1.92x the blocking review threads, 1.39x the comments, 1.24x the reviewer iterations, 1.19x the time to merge. Build failures run ~1.3x and sanitizer findings ~1.3x above the human cohort, which is Fung's CI/build jam with a ratio attached. And AI changes are structurally what this page describes: median 89 lines changed against 33, 3 files touched against 2.
Where it conflicts. This page's most alarming numbers are post-deployment — bugs per developer +54%, incidents per PR +242.7%. Google's post-deployment signal runs the other way: revert rate ~0.9x, below parity. Once AI code clears review and presubmit checks it is less likely to be rolled back than human code. Its Correctness-and-Safety static findings are also below parity (0.94x), as are lifetime/ownership hazards. The excess is concentrated in efficiency and coupling, not in breakage.
(Nothing above supersedes Faros's figures — they are different measurements, and both stand as reported. What follows is the weighting.)
How the two are reconciled, and what is left genuinely contested. Four axes of non-comparability absorb most of the gap:
- Different dependent variable. A revert is a specific remediation action; an incident is a production event. A codebase can generate more incidents and fewer reverts if the incidents are handled by forward fixes — which is the normal monorepo practice Google describes.
- No control cohort on this side. Faros compares low-adoption to high-adoption quarters within a company; Google compares AI-authored to human-authored code within the same window. Faros's design cannot separate "AI code is worse" from "orgs that adopt AI hardest are also changing in other ways"; Google's can, and does.
- Population maturity is exactly the disputed variable. Google's monorepo has centralized review, mature static analysis, and presubmit gates, and the paper's own reading of its split result is that the gates catch the fatal errors. That is a direct data point for DORA's "strong foundations protect you" and against this page's maturity-independence claim — from telemetry rather than survey, which was the axis Faros's critique of DORA rested on.
- Magnitude, where the sign agrees. Faros's median time-in-PR-review is +441.5%; Google's time to merge is +19%. Both above parity, an order of magnitude apart. Population and unit differ enough that neither refutes the other, but a reader carrying the +441.5% figure into an enterprise-monorepo context should carry the 1.19x alongside it.
Weighting. Google is the higher tier (empirical, controlled, production) and it is the more direct test of "does AI-authored code break more." But its COI runs opposite to Faros's and is just as directional: Google engineers measuring the output of Google's own AI coding tools, publishing a discussion that attributes the observed weaknesses to "historical default system configurations" rather than to the models. Two vendor incentives pointing in opposite directions, one of them attached to a control cohort. The honest verdict: the quality-down claim survives on review burden and pre-merge instability, and does not survive as stated on post-deployment stability in a mature review-gated org.
A second vendor panel, one quarter later — same shape, and the saved hours go missing (2026-08)#
DX's State of AI Impact in Engineering: Q2 2026 (Justin Reock, vendor-claim, 500+ customer organizations) is the closest structural match to this report the vault holds: an engineering-metrics vendor reading its own customer base, published as a newsletter readout of a gated PDF, with the prescription pointing at the vendor's product category. It is a second commercially interested instrument in the same market, not an independent replication, and every figure below is a summary number whose methodology the vault does not have.
Corroboration on the size driver, from a third kind of contrast. DX reports median PR size nearly doubled between Q1 and Q2 2026, and reads rising PR size as an early technical-debt indicator for exactly the reasons this page does — more complexity per change, more to review. Three sources now agree on direction while measuring three different things:
| Source | Contrast | Result |
|---|---|---|
| Faros (this page) | low- vs high-AI-adoption, within company | avg PR size +51.3%, files/PR +59.7% |
| Tran et al. | AI- vs human-authored changes, same monorepo | median 89 lines vs 33, 3 files vs 2 |
DX (vendor-claim) | Q1 2026 vs Q2 2026, calendar time in-panel | median PR size ~2x |
DX's is the only one measured on the calendar rather than across a cross-section, and it is the fastest-moving: a doubling of the median in two quarters. That makes this page's first open question (do the quality effects survive normalization for PR size?) more urgent rather than answering any part of it — DX publishes no size-normalized outcome at all.
The new finding, and it is a budget claim rather than a quality claim. DX estimates AI users save 4-6 hours per week, and reports that the innovation ratio — the share of time spent building new features versus maintenance and overhead — is flat over the same period. So the hours are real and the portfolio did not move.
This page and that finding compose into a hypothesis neither source establishes: the saved hours are being consumed by the downstream work the same throughput creates. Larger PRs to review, longer queues, more incidents, more rework — every quantity on this page is denominated in engineer-hours, and a flat innovation ratio is what it looks like when a velocity gain is spent paying for itself. DX supplies no decomposition of where the hours go, so this is a reading of two vendor datasets and not a measurement. It is also the org-layer answer to Ng's promotion story — the QA burden falling is supposed to free attention upward, and at panel scale the freed attention has not landed on new features.
DX ran the regression on its own composition, and it mostly does not close (2026-09-22). AI accelerates output, not innovation (Grace Fu, DX research analyst, 2026-09-09, vendor-claim; a follow-up that builds explicitly on the Q2 report, on 500+ DX customers — DX does not say whether it is the identical panel) updates the hours and then regresses the innovation ratio on 15 workflow metrics. The hours doubled on the calendar: self-reported time saved 3.0 h/week in Q3 2025 to 6.1 h/week in Q2 2026, and an "AI output" composite (AI-authored code share, agent-delivered work, PR throughput) explains 63% of the variance in time saved. Against the innovation ratio the same composite carries a standardized β of 0.16 (p < 0.01); the largest coefficient in the model is information-seeking friction at β −0.19 (p < 0.01) — time lost hunting for context, documentation, or answers — with deploy frequency (+0.06) and meeting-heavy days (+0.04) the only other coefficients DX shows, and those two only in the chart; the full 15-metric model explains 13% of innovation-ratio variance. DX's own reading: no single metric reliably predicts the ratio, treat innovation as a separately managed objective rather than a by-product of AI strategy, and re-run the regression on your own data rather than trusting the aggregate coefficients.
Three things this does and does not do for the hypothesis above. It confirms the shape — the input that predicts saved hours strongly predicts the portfolio shift barely, which is what the hours-spent-paying-for-themselves reading requires, now on one vendor's own instrument rather than on a juxtaposition of two vendors. It does not supply the decomposition the paragraph above asks for, and structurally cannot: the predictors are workflow metrics, not where-the-hours-went categories, so an 87% unexplained residual is compatible with the hours being eaten by review, QA and incidents (this page's mechanism) and equally with their going anywhere else — and the sink DX does name, information-seeking, is not on this page's list of downstream costs; it is closer to the context-starvation Faros's successor reads out of its restarts figure. And the number that looks most like support is the one to discount hardest: AI output → time saved at 63% is a predictor whose code-share component the Q2 report carried as self-report (which components of the regression's composite are telemetry is unstated) against a self-reported outcome from the same respondents, so common-method variance is the first candidate explanation for the strong fit, while the innovation ratio — a different self-report sharing no such channel with the workflow metrics — is exactly the outcome the model cannot explain (see Telemetry vs. Survey Measurement). Read the pair as: the instrument that asks developers about themselves agrees with itself strongly, and with the effort share it also asked them about weakly. The composition on this page stays a reading rather than a measurement, and DX's R² is roughly the size of the gap.
And a perception measure moving the way this page predicts. DX's Developer Experience Index fell 67 to 65 over four quarters, with the striking cut being a divergence between two of its component measures since Q1 2026: Code Maintainability +3.8% while Change Confidence -6.1%. DX's framing is that two historically correlated metrics have come apart — AI makes the code in front of you easier to understand while making what you push harder to trust.
Read the instrument before the finding. By the article's own definitions these are perceptions — maintainability is "how easily developers can understand the codebase," change confidence is "their trust that modifications won't cause production failures" — so this is the survey half of DX's instrument, not telemetry. That matters in this page's favour rather than against it: the perception-lags-reality argument predicts exactly this ordering, with felt confidence eroding a few quarters after the system outcomes Faros measured. A falling confidence index during a period of rising throughput is what perception catching up looks like. What it is not is independent confirmation of the incident and bug numbers, which remain Faros's alone.
DX's sixth finding is the budget frame around all of it: median quarterly organizational AI spend rose ~$1.5K to ~$44K over four quarters with tech-sector spend up nearly 28x, and its warning is that leaders who cannot connect that to feature velocity, innovation ratio or quality "may face increasingly difficult budget conversations." See Firm AI-Spend Intensity and Headcount Growth for where that sits among the vault's other spend instruments.
The whiplash as one team told it, with a mechanism attached (September 2026)#
Every number on this page is pooled across customers. Stolze & Strässle (ESEM 2026 SEIP, case-study) supply the corpus's first first-hand account of the failure mode at the level of one project, from an engineering manager at an energy-utility frontend team (P2). AI-generated changes accumulated over several months without proportionate review capacity; once the architectural problems became visible, a substantive feature had to be discarded and reimplemented from scratch. His own counterfactual is the mechanism this page argues for, stated by the person it happened to:
"had this been done without an AI system, we would not have generated so much code. .. maybe we would have noticed earlier"
Two things make this worth carrying despite being n = 1 inside an n = 5 study. It names the detection-latency channel specifically — not that AI code is worse, but that volume postpones the moment a problem becomes visible, so the same defect surfaces later and costs a rewrite instead of a fix. And the authors resist the obvious reading: P2 "framed this as a failure mode rather than an inherent property," with the lesson that "productivity gains and validation infrastructure need to be scaled together — an organizational choice rather than a purely technical optimization." That is the same conclusion this page reaches from telemetry, reached from the inside of one project. The whole paper's answer to the pressure is structural rather than corrective — distribute supervision across three layers instead of scaling review (Layered Supervision).
The second human-controlled measurement, and it is outside the monorepo (September 2026)#
Kraishan (arXiv 2609.17598) is the first source to test this page's post-deployment claim in the population Faros's customers actually resemble — public GitHub repositories with weaker and heterogeneous gates, not Google's centrally-reviewed monorepo. 37,623 provenance-labelled PRs, 2,807 repos, and the control this page has been asking for twice over: 4,027 human PRs kept only from the 810 repositories that also contain agent PRs, same window, capped per repository.
It lands on Google's side, harder. Within 90 days of merge, human PRs are reverted at 11.5%; Codex PRs at 6.1% (OR 0.50), Devin at 14.5% (OR 1.31), and Copilot, Cursor and Claude Code are statistically indistinguishable from humans after BH correction. Weighting Table 2's rows gives a pooled agent rate of ≈7.7%, OR ≈0.64 (arithmetic from the table, not a figure the paper reports). So the sub-parity revert result does not depend on Google's presubmit gates — it survives the population change, and it survives it by a wider margin. Per-line post-merge churn goes the same way: four of the five vendors sit at or below the human level, with Claude Code lowest of any group (δ = −.33, median 0.5 follow-up commits per 100 changed lines against the human 5.3).
Three reasons this does not simply refute Faros, and they are the same three as before. The dependent variable is a revert flagged by commit message on the PR's most-changed file, which the authors say misses silent rewrites — an incident handled by a forward fix is invisible to it, exactly as it was in Google's data. The population is >100-star open source, three languages, and PRs that cleared a public project's review; Faros's is enterprise SDLC telemetry across whole orgs. And the authors offer a mechanism that is not about the code at all: "the humans steering them may assign bounded, well-specified tasks" — a selection effect that would produce sub-parity reverts from an agent of any quality.
What genuinely moves. Two empirical sources with human control cohorts, in two populations with opposite gate maturity, now both find agent code at or below parity on post-merge failure, while this page's vendor telemetry reports incidents/PR +243%. The quality-down claim is now cornered into the pre-merge and review-burden half of the pipeline on every controlled measurement the vault holds. (Nothing above supersedes Faros's figures; they remain as reported, on a different dependent variable and a different population.)
And the new variable this page cannot see at all. Faros pools every AI tool into "AI adoption." Kraishan's whole result is that the spread between vendors is wider than any agent-vs-human gap — 0.50 to 1.31 in revert odds against the same baseline, a sixfold spread in security-smell presence, a fourfold spread on credential smells. An adoption-depth cross-section over a customer base with an unknown and shifting tool mix is averaging over that spread, which is a source of variance the report never accounts for and a plausible contributor to the "maturity doesn't protect you" pattern: a cohort that adopted Devin and a cohort that adopted Codex are not the same experiment.
Faros's own successor: the shock eases, the strain moves downstream (September 2026)#
The Speed Trap (Faros Research, 2026-09-18, vendor-claim) is the sequel to this page's source — the Q3 2026 AI Engineering Report, same 22,000-developer / 4,000-team panel, now over the most recent 12 months. It is the only source in the corpus that re-runs this page's own instrument, and it is the same vendor with the same commercial interest, so it is a second reading, not a replication.
Read the construct before any number. Every percentage in the post is a period-over-period delta — this dataset against the prior (~April 2026) report's — and both windows are already-high-AI-adoption teams. There is no AI-vs-non-AI cohort, no definition of what counts as an "AI-assisted" change, no sampling description, no significance testing, and the full report is behind a lead-gen form. Where the prior report compared a company's low-adoption quarters to its high-adoption quarters, this one compares deep adoption to deeper adoption. So it cannot answer any question on this page that needs a control cohort, and the figures below are changes in a rate of change, not effects of AI.
The deltas, side by side. Rows transcribed from the report's Q3 2026 "Key Findings" infographic (viewed directly; the body text states only some of them). Nothing here supersedes this page's figures above — those are the prior window's deltas and remain true of it.
| Metric | Q3 2026 report | Prior (~April 2026) report |
|---|---|---|
| Deployments per week | +13.8% | −11.7% |
| Code deletion ratio | +71.6% | +861% |
| Incidents per PR | +14.5% | +242.7% |
| PRs skip review entirely | +76.3% (see caveat) | +31.3% |
| Work restarts per developer | +66.7% | +13.8% |
| Average task time in QA | +300.6% | +33.7% |
| Monthly incidents | +125.4% | +57.6% (this page, from the prior report itself, reads +57.9%) |
Body text adds one metric the chart does not: average PR size +71.8%, "even more than in our previous dataset" (which was +51.3%). It is a different statistic from the chart's code-deletion ratio +71.6% despite the near-identical number — do not conflate them. And the +57.6%/+57.9% cell is the only place two Faros publications disagree about one Faros number; the gap is trivial in magnitude but it is the reader's only available check on the fidelity of that whole right-hand column.
Construct caveat on the no-review figure — carry it wherever the number goes. The prose reads "six months ago… a 31.3% rise; that figure has increased to 76.3%." The infographic renders it as "+76.3% PRs skip review entirely, up from +31.3% in prior dataset," formatted identically to every other growth row on the chart. On that evidence the vault reads 76.3% as a period-over-period growth rate in the unreviewed-PR metric, not as the share of PRs merged without review. The prose alone is genuinely ambiguous and the definitions live in the gated report. Do not state that 76.3% of PRs are merged without review.
Adoption, and the frontier moving to agents. 79% of developers use at least one AI tool weekly; 86% of teams exceed 50% weekly-active-user penetration; AI-code acceptance 65% (against the 60% on AI as Primary Author from the prior report). Agentic review now runs on 50–80% of PRs at many companies (prior report: 0%→25%), and agents are opening 13–14% of PRs at a leading edge — against "<1% of PRs" in the prior dataset. Those two authoring figures have different denominators (a pooled panel share versus an unspecified leading-edge slice), so they do not compose into "the panel crossed from <1% to double digits," which matters for the predictions parked on this page's neighbours.
What the shape actually says. All three "initial shock" metrics decelerated hard; all four downstream metrics accelerated. Faros's own reconciliation is that per-change quality stopped collapsing while volume kept rising, so total operational cost keeps climbing: incidents per PR +14.5% against monthly incidents +125.4%. That is a genuine refinement of this page's thesis rather than a retraction — the whiplash relocates from the change to the system — and it is the first reading in which the scariest number on this page (incidents/PR +242.7%) is described by its own author as a transient of the adoption transition.
The normalization gap is not closed, and the successor makes it more conspicuous. The post never relates incidents per PR to PR size, anywhere. What it does supply is that the two moved in opposite directions across the window: per-change incident growth collapsed 242.7% → 14.5% while PR-size growth accelerated 51.3% → 71.8%. If larger PRs were driving the incident rise, those two should track. That is informative and it is not a normalization — two uncontrolled period-over-period deltas with no joint model and no size-stratified incident figure. See this page's first open question.
The thrash changed shape, and the prescription did not. The prior report's dominant strain was parallelism (too many concurrent PR contexts); that is easing, and restarts — throwing in-progress work away and re-approaching from scratch — take its place at +66.7%, "nearly five times" the prior +13.8%. Faros reads this as agents lacking sufficient context to get the work right the first time, which is the same diagnosis pointing at the same remedy as recommendation #10 in the prior report (the context engine). The failure mode moved; the product category it implies did not, which is worth noting before reading the mechanism as a finding.
Takeaway 8 is the one prescriptive claim, and it carries no number. Teams with heavy agentic-review adoption see "faster first reviews and lower change-failure rates" — no percentages for either, explicitly flagged correlational by Faros, and with the concession that unreviewed merges keep rising even where agentic-review adoption is high. It is the first outcome-flavoured signal in the corpus that agentic review associates with a quality outcome, and at this evidence weight it does not touch the contested edge on Review as the Control Point. Faros still concludes the leverage is at authoring, with review as the second line of defence — unchanged from this page's thesis.
One thing cuts mildly in the vendor's favour. A vendor whose narrative is "the crisis deepens" reported its own headline alarm metric decelerating by a factor of roughly seventeen. That is the direction that costs the seller something, which is weak evidence against the pure-incentive reading of this page's third open question. The offsetting observation: the metric that now leads the alarm — QA time +300.6% — is the figure in the whole report with the least corroboration anywhere else in the corpus.
Connections#
- Layered Supervision — the field response to this pressure and a qualitative instance of the failure: a team that let AI-generated change accumulate past its review capacity and lost a feature to a rewrite. Its own framing is that the fix is neither more review nor better authoring but a redistribution of the supervision function across preventive, executable and human layers — and its refusal of the word "maturity" cuts against this page's maturity-independence claim from a third direction
- Community Smells Under AI Adoption — the same org-scale question from survey rather than telemetry, and with the opposite sign: self-reported team social health improves under AI adoption where Faros measures quality degrading. Different constructs and different instruments rather than a direct contradiction — but the study's own free-text minority sounds much more like the telemetry than like its coefficients
- AI as Primary Author — the precondition: the assistant→author threshold (60% code acceptance) is what floods the system; "AI is the primary author now" is the report's framing of who generates the absorbed output
- Systems Thinking Over Specialization — a named counter-strategy: Netflix's paved paths and design systems try to raise absorption capacity through infrastructure-encoded guardrails rather than process gates
- Excellence as an Operating System — the cultural bet this page's data challenges: Stone's "adding process never got better outcomes" vs. Faros's evidence that even high-maturity orgs fail to absorb agent-scale throughput; whether talent density + encoded guardrails substitutes for process is the open question recorded on both pages
- Verification as the New Bottleneck — Faros corroborates the bottleneck with telemetry but refines the fix: improve authoring, don't just scale review. That page now also carries CircleCI's Q2 2026 Pulse (
vendor-claim, 20M+ CI workflows), a second independent vendor telemetry set pointing the same way: feature-branch throughput +7.7% YoY against flat main-branch throughput, and an elite/median velocity gap that widened 8× → 9× in one quarter — the widening-gap shape this page argues for, measured on a different instrument by a vendor with a different product to sell. One reading cuts against this page though: CircleCI's main-branch success rate improved 70.8% → 76.7% over the same window, a quality metric moving the right way. Different layer (CI pass rate vs production incidents) and a CircleCI-customers-only population, so it is a boundary marker rather than a refutation — the pipeline can get greener while what it ships gets worse - AI Brain Fry — the cognitive-load channel: context-switching and under-review are the human-side strain the whiplash induces at org scale
- Psychological Costs of AI Adoption — the same review load, counted in people instead of PRs. Faros measures median time-in-PR-review up 441.5% and cannot say who absorbed it or what it cost them; that case study interviews the absorbers one year in, at a firm where the same shift is underway, and reports what the telemetry has no column for — the review burden is levied by retained accountability, not by output volume ("if we weren't responsible for the code it produced, it would be a lot faster"), it is unequally distributed by role (software engineers voiced accountability anxiety 10/10, architects 0/3), and the organization measured none of it while measuring adoption progress closely. It also supplies a mechanism for why the load does not fall as models improve: delegating authorship converts verification from something interleaved with writing into a separate post-hoc phase that must rebuild the understanding authorship used to supply for free
- Agentic Technical Debt — the quality degradation is debt compounding industrially; Faros's "context engine" (rec #10) is the CLAUDE.md mechanism scaled to the org
- Telemetry vs. Survey Measurement — the methodological basis for the maturity-independence claim and the explicit DORA counterpoint
- Vibe Coding vs. Agentic Engineering — the dark mirror: this is what the data looks like when orgs fail to preserve Karpathy's quality bar
- Blast Radius (Agentic) — Faros's "wider blast radius per change" is the code-change sense (larger PRs reaching further into the codebase), distinct from the security-compromise sense
- Harness Shrinkage as Models Improve — a counter-pressure data point: even as models improve, org-level quality degrades, because the harness (context provisioning, quality gates) didn't keep pace with capability
- Outsource Your Thinking, Not Your Understanding — the senior-engineer tax is comprehension debt cashed in at review time: someone must reconstruct the intent the author never held
- Agentic Coding Work-Composition Shift — the juxtaposed telemetry: Anthropic's same-period study finds session value up ~27% and debugging down at the interactive-session layer, while this finds quality down at the org-SDLC layer — different units (session success vs downstream incidents), and
empiricalresearch telemetry vs thisvendor-claimreport - Organizational Complements to AI — the diagnosis under the whiplash: throughput up but quality down is the productivity-paradox failure mode when orgs adopt AI faster than they redesign the review/QA complements; the missing complement is what the telemetry catches
- AI Investment Story, Not Efficiency Story — the financial-metrics sibling: throughput up but realized per-head efficiency lags because the org complements lag adoption — the same lag as the whiplash, measured on revenue-per-employee instead of SDLC quality
- The Three Loops of AI-Native Building — the telemetry that outranks Andrew Ng's self-report: he claims self-testing agents cut the developer's QA burden "significantly," where median time-in-PR-review rose 441.5%; scope (0-to-1 personal builds vs production orgs) is the likely reconciler
- Security Debt of Agent-Generated Code — non-vendor
empiricalcorroboration of the quality half on a security axis (38.9% of agentic PRs carry a security smell, concentrated in CI/container files), and the source of the PR-size-stratified data this page's first open question asks for - Agent-Generated Test Quality — the authoring-quality thesis measured inside the test suite (
empirical, non-vendor): agent-authored tests carry unmocked file I/O and non-determinism at ~1.4× the human rate, one concrete channel from throughput rise to the CI/build jam. It also complicates this page's dismissal of Ng — self-testing agents do broaden coverage (edge-case variety 0.62 vs 0.32), they just destabilize the runner, so both sides of that dispute get a piece - Risk-Tiered Auto-Approval — the risk-tiered gating this report recommends, running in production: PostHog's StampHog auto-approved ~1 in 3 PRs merged into their main repo behind a deny-list + a <500-line/<20-file ceiling. It is a review-layer lever, which is what this page argues is the wrong end — but its size ceiling attacks the "wider blast radius per change" driver directly, and the paired practice (decompose into stacked sub-400-line PRs) is an authoring-side change
- Agent-Vendor Heterogeneity — the second human-controlled
empiricalsource to land against this page's post-deployment half, and the first outside an enterprise monorepo. Pooled agent reverts ≈7.7% against a same-repo human 11.5%, per-line churn at or below human for four of five vendors — in >100-star public GitHub repos, which removes the "Google's gates caught it" explanation that absorbed the Tran result. It also names a variance source this page's pooled telemetry structurally cannot see: revert odds run 0.50 to 1.31 between vendors against one baseline, so an adoption-depth cross-section over customers with an unknown tool mix is averaging across a spread wider than the effect it reports. See the September 2026 section above - Efficiency Debt of AI-Generated Code — the human-controlled foil, and the only source that measures both halves of this page in one dataset (see the section above). It corroborates the review-burden half at a smaller magnitude, contradicts the post-deployment half (revert rate ~0.9x), and adds a downstream cost this page never counted: ~5% relative compute and ~8% relative memory overhead for AI-heavy functions, driven by explicit loops replacing standard-library calls. Its
#oqabout weaker presubmit gates is the mirror of this page's maturity question - Agent Review Comment Resolution — whether the fastest-automating layer in this report's data actually lands. Faros records agentic review going 0% to 25% of PRs while agentic authoring stays under 1%; that study measures the output of that layer across 54,713 comments and finds roughly seven in ten resolved (Copilot 72.9%, Cursor 67.2%, Codex 54.8%). A mild counterweight to the whiplash framing — the automated oversight layer is being used, not ignored — with two limits: the pooled figure is 83.5% Copilot, and nothing there connects comment adoption to any downstream quality outcome, so it does not touch this page's incident and bug numbers
- The Committed-Artifact Chain — the same diagnosis, asserted rather than measured, plus the fix this page argues against. Anthropic's Applied AI playbook (
vendor-claim, 2026-08-21) opens on exactly this finding — build collapses to hours while the stages either side stay human-paced, controls stop matching reality, governance cost rises — and then prescribes an artifact chain whose bulk is review-and-gate machinery, which is the end this page's thesis calls the wrong one. Two of its eleven plays are authoring-side (CLAUDE.md/skills constraining generation; a feedback loop that makes the session verify before a human sees it), so the disagreement is one of proportion rather than principle. Where the two collide directly is review time: the playbook expects time-to-first-review to "fall to minutes" and review time per PR to fall once tests catch what reviewers used to, against the +441.5% median time-in-PR-review measured here. Those are different quantities and can both move — an agent reviewer posts in minutes without shortening the human's read — but only one side has measured anything, and it is this one - Review as the Control Point — the non-vendor foil. Faros (
vendor-claim) argues the gap widens with adoption and maturity doesn't protect; this CMU theory (empirical, non-vendor) argues AI does not fix the sign of the effect — the team's expertise + process do — and its own GitHub telemetry finds the agent no-review rate converging down toward the human baseline over mid-2025→early-2026, not widening. Caveat both ways (open-source vs enterprise populations; neither out-measures the other on quality outcomes) - The Code-Quality Payoff Is Token-Indexed — the argument that would justify accepting this quality drop, and what it actually rests on: DHH holds that architectural coherence is worth paying for only while tokens are scarce
- Open Source Under Agent Contributions — the single-maintainer version of unabsorbed output: an open-source PR backlog doubling weekly, answered by delegating review rather than by staffing it
- Does the Augmentation/Automation Split Govern Skill at Work? — the 31.3%-of-PRs-merged-unreviewed figure used as the workplace's third use mode: automation-mode users in the randomized learning experiment still route output through themselves, while abdication removes the review step that would have been the learning — a failure mode generated by volume, which a proctored session holds at one
- Experimental Learning Impact of Generative AI — the learning-side boundary on this page's absorption story: the randomized experiment's automation arm still routes output through the human (gains hollow once AI is removed, but real in the short run), whereas 31.3% of PRs merged with no review is the review step vanishing outright — output unabsorbed, and the learning that absorbing it would have produced not happening either
Open Questions#
- Faros's own deferred question: do the bug/incident increases persist when normalized for PR size, or do larger PRs account for most of the quality deterioration? (If the latter, hard PR-size limits are the highest-leverage fix.) Partially answered by Security Debt of Agent-Generated Code (
empirical, non-vendor): on the security axis, PR-level flagging rises monotonically with change size — 16.2% for 1–9-line PRs to 53.6% for 1000+-line PRs, a 37.4-point spread — which is the size-stratified evidence this question asks for and supports hard PR-size limits as a real lever. Two gaps keep it open: it measures smells introduced, not the bugs and incidents Faros counts, and being a cross-sectional association it can't say whether capping size lowers density or merely re-partitions the same changes across more PRs. Further partially answered 2026-08-12 by Tran et al. (empirical, with a human control cohort): AI changes there are indeed larger (median 89 lines vs 33, 3 files vs 2), and the downstream comparisons are stratified on change size among other covariates — so the ratios that survive stratification are not the size effect. What survives is split by outcome: blocking threads 1.92x and build failures ~1.3x stay above parity, revert rate ~0.9x stays below. So size does not account for the deterioration, and the deterioration does not have one sign. A hard PR-size limit therefore addresses review burden rather than production stability, which is a narrower case for the lever than this bullet originally assumed. A third partial, 2026-09-22, on a public-GitHub population: Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild (Kraishan,empirical, non-vendor, same-repo human control) runs the size question two ways and both say size is not the story. Its size-stratified sensitivity check finds the pooled agent-human security-density difference concentrated entirely in the XL PR bucket (δ = −.08, p =.025) with no difference in the four smaller buckets — which is the opposite reading of the same shape the security study gives: there, larger PRs carry more smells; here, larger PRs are where agents beat humans. And its post-merge churn measure is size-normalized by construction (follow-up commits ÷ changed lines), with four of five vendors at or below the human level after normalization; the unnormalized ranking is different (Copilot lowest of any group, Codex slightly above humans), so normalization changes who wins but does not create an agent penalty at either setting. What it still cannot supply is the quantity this bullet actually wants: it counts reverts and churn, not the bugs and incidents Faros counts, and its revert detector works by commit message and misses silent rewrites. The lever is now supported by two sources on review burden and by none on production stability. Faros's own successor was asked this question and did not answer it (2026-09-22): The Speed Trap re-runs the same panel six months on and still relates incidents per PR to PR size nowhere in the post — no size-stratified incident figure, no normalization statement, nothing. What it does supply is a suggestive non-answer: across the two windows the per-change incident growth collapsed from +242.7% to +14.5% while PR-size growth accelerated from +51.3% to +71.8%, so the two quantities this bullet asks to be related moved in opposite directions. If size were driving the incident rise they should track. That is evidence against the strong form of the size hypothesis on Faros's own instrument, and it is emphatically not a normalization: two uncontrolled period-over-period deltas on an already-high-adoption panel, no joint model, no control cohort, and the methodology is in a form-gated report. The trigger is unchanged — incidents per PR cut by PR-size bucket, from anyone. - Code churn +861% is genuinely ambiguous (Faros lists three explanations: rework of AI code, productive legacy refactoring, or accelerated polish). The cross-customer metric can't resolve it — a real gap, not a finding. A partial proxy added 2026-07-29 by 5 takeaways from the State of Software Delivery Q2 Pulse report (
vendor-claim): CircleCI's Merge Efficiency Ratio — validation cycles a feature branch needs before it lands on main (median 3.9, top-5% 2.6, elite cohort 1.3) — counts a pre-merge form of the same rework, and it is countable per team rather than pooled cross-customer. It narrows the ambiguity from one side only: cycles spent failing validation before merge are hard to read as "productive legacy refactoring," so a high MER is closer to unambiguous rework than churn is. It does not decompose Faros's metric, because the two measure different things — MER counts attempts, churn counts lines-deleted-to-added, and a clean refactor that passes CI first try is invisible to MER while dominating churn. A second partial from the other side of merge, 2026-09-22: Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild measures post-merge churn with a same-repo human control — follow-up commits on the PR's most-changed file within 90 days, divided by changed lines. Four of the five agents sit at or below the human level (Claude Code δ = −.33, median 0.5 against the human 5.3; Copilot −.16; Cursor −.08; Codex and Devin at parity). That does not decompose +861% either, and the constructs are again different — commits-per-changed-line over 90 days against lines-deleted-to-added within a short window, and this one is measured on one file per PR as a cost-driven proxy for whole-PR churn. What it does is eliminate the first of Faros's three explanations in this population: if the churn were rework of AI code, agent-authored changes should draw more subsequent commits per line than human-authored ones in the same repositories, and they draw fewer. The remaining two explanations (productive legacy refactoring, accelerated polish) are both about work that is not attributable to a specific PR, which is precisely what a per-PR design cannot see. The metric was re-measured by its own publisher and it does not discriminate either (2026-09-22): The Speed Trap reports the same quantity — now named "code deletion ratio" rather than "code churn" — growing +71.6% against the prior window's +861%. A growth rate that falls by more than an order of magnitude is consistent with all three of Faros's explanations (transition rework subsiding as teams learn the tools; a legacy-refactoring backlog being worked off; polish normalizing), so it discriminates among none of them, and the report offers no decomposition. Two things it does settle, both narrow. The +861% was a transient of the adoption transition, not a new level — the runaway-and-compounding reading of this bullet is out. And Faros uses two different names for what is presumably one metric across two reports with no definition published for either, which is an independent reason the construct cannot be decomposed from the outside. What would settle it is unchanged: a per-team decomposition of deleted lines by whether the deleted code was itself recently AI-authored. - How much of the "maturity doesn't protect" claim survives the vendor incentive to argue exactly that (i.e., "your existing practices won't save you — you need our platform")? Partially answered by Review as the Control Point (non-vendor,
empirical): its whole thesis is the opposite — AI doesn't fix the sign; team expertise and process do — which leans toward DORA's "foundations protect you" and against Faros's determinism. But it argues the moderators exist rather than measuring a maturity effect, so the vendor-incentive question isn't closed, only counterweighted by a non-vendor source that disagrees with the framing. A third position, 2026-09-22 — When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering (case-study, 5 interviews, non-vendor, academic) refuses the axis rather than taking a side on it. Its five teams varied along three named dimensions — governance posture, system criticality and homogeneity, and team composition — and the authors decline the word maturity explicitly, "since it suggests a linear progression that our data does not support." That is a direct challenge to the shared premise under both this claim and DORA's, which is that organizations can be ordered on one scale at all; if the space is genuinely multi-dimensional, "maturity doesn't protect you" and "foundations protect you" are both answers to a malformed question. It does not close the vendor-incentive question and cannot: refusing an ordering on five qualitative points is a modelling choice, not a measurement, and the survey behind those dimensions was convenience-sampled and unpiloted. What it does is put a non-vendor source on record that the ladder is the thing in doubt. A weak datum from the incentive's own side, 2026-09-22: in The Speed Trap Faros reports its own headline alarm metric — incidents per PR — decelerating from +242.7% to +14.5%, and states plainly that per-change quality "is no longer collapsing." A vendor selling on "your practices won't save you" publishing a seventeen-fold softening of its scariest number is the direction that costs the seller something, which is mild evidence against reading the maturity claim as pure narrative construction. It is mild: the alarm did not go away, it moved to metrics the same product addresses (QA time +300.6%, monthly incidents +125.4%, unreviewed merges), and the successor is the same instrument with the same undisclosed method, so it cannot audit the prior report's framing from outside. The question still needs a non-vendor measurement of a maturity effect, which nothing in the corpus supplies.
Resolved Questions#
- Faros reads under-review as a widening crisis; CMU's non-vendor GitHub telemetry finds the agent no-review rate converging toward the human baseline (>50%→~14%) as orgs learn to review agent code. Is the divergence real (enterprise vs open-source populations, adoption-depth cross-section vs calendar-time trend) or does the whiplash's under-review pressure only surface where PR volume is highest? Answered: The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence — mostly not real: Faros's +31.3% is a delta in unreviewed-PR count across adoption depth (enterprise, all PRs) while CMU's is a falling share of unreviewed agent PRs over calendar time (open source) — a falling rate and a rising count coexist under Faros's own volume growth. The volume clause is supported (median per-project no-review ≈0% vs pooled >50%; triage by PR type), and Faros's own risk-tiered-gating remediation is the triage behavior CMU observes emerging. The residual disagreement is a forecast: does triage discipline survive agentic authoring crossing from <1% to double digits — untested in both datasets.
Sources#
- When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering — Stolze & Strässle (OST Eastern Switzerland UAS / smartive AG, arXiv 2608.26316, 2026-08-26, ESEM 2026 SEIP),
case-study: §5.5 (P2's discarded-and-reimplemented feature and the scale-them-together lesson) and §4.5 (the three configuration dimensions and the explicit refusal of "maturity"). Five interviews, one convenience-sampled survey, no measured outcome — evidence notes at Layered Supervision - Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild — Obada Kraishan (Texas Tech), arXiv 2609.17598, 2026-09-12,
empirical, non-vendor, sole author with no venue: §3.1 (the matched 810-repository human baseline), §4.1 (the XL-bucket size-stratified sensitivity check), §4.3 + Table 2 (size-normalized post-merge churn and 90-day revert outcomes), §5 (the two mechanisms the design cannot separate), §6 (revert-by-commit-message misses silent rewrites; agents self-select to tasks). Tables 1 and 2 verified cell-for-cell againstpdftotext -layout. Full treatment and the provenance caveat at Agent-Vendor Heterogeneity - AI Engineering Report 2026: The Acceleration Whiplash — Executive Summary, Findings #1–7, "What Engineering Organizations Should Do"
- The Speed Trap: 8 takeaways from our latest AI engineering research — The Speed Trap: 8 takeaways from our latest AI engineering research, Faros Research (corporate byline), 2026-09-18,
vendor-claim— the successor to the row above, same vendor, same 22,000-developer / 4,000-team panel. All eight takeaways plus the "AI Engineering Report Q3 2026 | Key Findings" infographic, viewed directly and transcribed in the raw; four of the seven charted deltas (deployments/week, code deletion ratio, and both prior-report columns for incidents/PR and monthly incidents) appear only in the image. Construct warning attached to every figure: each percentage is a period-over-period delta between this dataset and the ~April 2026 report's, both on already-high-adoption teams — no AI-vs-non-AI cohort, no definition of "AI-assisted", no sampling description, no significance testing, and the full report ("10 recommendations") is behind a HubSpot form that was not filled. The "76.3%" is read as a growth rate in the unreviewed-PR metric, not a share of PRs — the chart formats it identically to every other growth row while the prose is ambiguous; see the caveat box above. Two further parse notes: the chart's "code deletion ratio +71.6%" is a different statistic from the body's "average PR size +71.8%", and the chart's prior-report column gives monthly incidents +57.6% where the prior report itself says +57.9% - 5 takeaways from the State of Software Delivery Q2 Pulse report — Jacob Schmitt, CircleCI blog (2026-07-08),
vendor-claim: findings 1, 2, 4 — the 8×→9× widening velocity gap, feature-vs-main throughput split, main-branch success rate 70.8%→76.7%, and the Merge Efficiency Ratio - Characterizing the Quality Profile of AI-Generated C++ in Production — Tran et al. (Google, arXiv 2608.06640, 2026-08-06),
empirical: §4.2 Table 2 (change structure), §4.3 and Figure 4 (revert / sanitizer / build-failure ratios and the five review-friction ratios), §6 (threats to validity). Full treatment, evidence note and COI at Efficiency Debt of AI-Generated Code - AI accelerates output, not innovation — Grace Fu (research analyst, DX), AI accelerates output, not innovation (DX Engineering Enablement newsletter, 2026-09-09;
vendor-claim, same tier and same COI as the Q2 report it builds on). Time saved 3.0 → 6.1 h/week (Q3 2025 → Q2 2026); AI output explains 63% of time-saved variance; multivariate regression of the innovation ratio on 15 workflow metrics over 500+ DX customers, standardized β: AI output +0.16 (p < 0.01), information-seeking −0.19 (p < 0.01), deploy frequency +0.06 and meeting-heavy days +0.04 (chart only,, viewed); model R² 0.13. The raw is a paraphrased digest, not a verbatim clipping — numbers and the model specification are preserved, prose is not quoted as DX's words. Whether it is the identical Q2 panel is not stated; no n per metric, no coefficient table beyond the four charted, no statement of which composite components are telemetry - The State of AI Impact in Engineering: Q2 2026 — Justin Reock, The State of AI Impact in Engineering: Q2 2026 (DX, Engineering Enablement newsletter, 2026-07-22). Evidence tier corrected
empiricaltovendor-claimat compile — a developer-productivity vendor's readout over a self-selected panel of its own 500+ customer organizations, published as lead-gen for a gated report and a webinar; full reasoning in the Sources entry. Findings 1-6: the 34%-to-52% self-reported code share, the near-doubled median PR size, DXI 67-to-65, the Code Maintainability / Change Confidence divergence, the 4-6 saved hours against a flat innovation ratio, and the$1.5K-to-$44K median quarterly spend. The methodology tables are in the gated PDF and are not in the vault, so no figure here has an n, a confidence interval, a stated measurement window per metric, or a statement of which quantities are survey and which are telemetry. Web article, no docling parse; charts on the page are images and every number quoted appears in the newsletter's own prose. The control-group claim is handled at Telemetry vs. Survey Measurement
Cited by 37
- AI as Primary Author×7
Faros's own successor moves all three numbers, six months on (2026-09-22). The Speed Trap (faros ai…
- Review as the Control Point×6
This page's standing forecast question names Faros as the predictor: does the no-review convergence…
- Telemetry vs. Survey Measurement×6
This is a flagged inter-source contradiction. DORA's 2025 State of AI-Assisted Software Development…
- Verification as the New Bottleneck×6
Discount appropriately. Both quantities are perception measures by DX's own definitions, from a…
- Open Questions Backlog×5
Acceleration Whiplash: Faros's own deferred question: do the bug/incident increases persist when…
- Risk-Tiered Auto-Approval×5
It supports the size gate, strongly. Security-smell prevalence in agentic PRs climbs monotonically…
- The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence×5
The arithmetic reconciliation is direct: a falling rate and a rising count coexist whenever volume…
- The Committed-Artifact Chain×4
Does an artifact chain make review deeper or only earlier? Faros measured review time exploding and…
- The Three Loops of AI-Native Building×4
The human didn't get removed from the loop; they got promoted out of QA. Notice this cuts against…
- Does the Augmentation/Automation Split Govern Skill at Work?×4
Acceleration Whiplash — Faros's org-scale telemetry across ~4,000 teams: daily PR contexts per…
- Efficiency Debt of AI-Generated Code×3
Two things follow. First, this is a production-scale null against the simplest reading of Review As…
- Excellence as an Operating System×3
Stone's anti-process stance sits directly against Acceleration Whiplash — Faros's telemetry showing…
- Faros AI×3
AI Engineering Report 2026: The Acceleration Whiplash — the paradox "sharpened into a crisis." See…
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?×3
Telemetry: 31.3% of PRs merged with no review, review time up ~5×, daily PR contexts per developer…
- Agent-Generated Test Quality×2
The reported gap is small — 88.08% vs 85.70% strong assertions — and the more interesting number is…
- Agent-Vendor Heterogeneity×2
Acceleration Whiplash — the direct counterweight on the post-deployment half. Faros reports…
- Agentic Coding Work-Composition Shift×2
Acceleration Whiplash — the juxtaposed telemetry: value/success up at the session layer here vs.…
- AI Brain Fry×2
Acceleration Whiplash — Faros Ai's org-scale telemetry of the same fatigue: daily PR contexts per…
- AI Investment Story, Not Efficiency Story×2
Acceleration Whiplash — the SDLC-telemetry sibling: throughput up but realized quality lags because…
- The Code-Quality Payoff Is Token-Indexed×2
No supersession. Both are unmeasured practitioner accounts from committed partisans, a week apart,…
- Open Questions Dashboard×2
Three Loops Of Ai Native Building: Ng asserts the developer's QA burden fell "significantly."…
- Psychological Costs of AI Adoption×2
Acceleration Whiplash — the telemetry and the interviews describe one phenomenon from two ends.…
- Security Debt of Agent-Generated Code×2
Acceleration Whiplash — non-vendor empirical corroboration of the quality half of the whiplash on a…
- Systems Thinking Over Specialization×2
She names the agent-scale endgame explicitly: Netflix's vision is "so many agents contributing to…
- Agent Review Comment Resolution
Acceleration Whiplash — Faros Ai records agentic review going 0% to 25% of PRs, faster than agentic…
- Agentic Technical Debt
Acceleration Whiplash — the same compounding mechanism measured at industry scale; Faros Ai's…
- Andrew Ng
QA was the job that went away. "Last year, a lot of developers (including me) were acting as the QA…
- Blast Radius (Agentic)
Acceleration Whiplash — different sense of "blast radius": Faros Ai's "wider blast radius per…
- Community Smells Under AI Adoption
Acceleration Whiplash — the same org-scale question answered from telemetry rather than survey, and…
- Experimental Learning Impact of Generative AI
Does the same use-mode split govern workplace skill accumulation (the open question Automation…
- Firm AI-Spend Intensity and Headcount Growth
Acceleration Whiplash — a third spend instrument, and the first that reports spend next to…
- Layered Supervision
Acceleration Whiplash — the pressure this is a response to, with a qualitative instance attached:…
- AI Coding Practice
Acceleration Whiplash — Faros 2026: AI floods a human-paced SDLC with output it can't absorb —…
- Open Source Under Agent Contributions
Acceleration Whiplash — the industry telemetry of the unabsorbed-output problem (review time 5×); a…
- Organizational Complements to AI
Acceleration Whiplash — the downstream-cost evidence of missing complements: when orgs adopt AI…
- Outsource Your Thinking, Not Your Understanding
Acceleration Whiplash — Faros Ai's senior-engineer "tax" is comprehension debt cashed in at review:…
- Vibe Coding vs. Agentic Engineering
Acceleration Whiplash — the dark mirror: Faros Ai's industry telemetry of what the quality bar does…
Related articles
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Review as the Control Point
Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Telemetry vs. Survey Measurement
Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (…
- Security Debt of Agent-Generated Code
Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…
