Sources#
- 2026 State of Scaling: The Great Sorting
- AI Engineering Report 2026: The Acceleration Whiplash
- AI-to-AI Code Reviews of GitHub Pull Requests
- Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild
- Thread by @AndrewYNg
The questions#
Six #oq/now items from the ai-coding-practice oversight cluster, answered as one synthesis. They share one structure: each asks whether an oversight number means what it appears to mean once the population behind it is taken apart.
- AI as Primary Author: what does "acceptance" mean when the agent applies the change directly and the human's acceptance is not reverting it?
- Closed-Loop AI Review: what is the AI-review rate on the 1.73M quarantined branch-only PRs?
- Closed-Loop AI Review: does the 58–65% same-product comment-volume gap survive holding the reviewer bot fixed and matching on change size?
- Agent-Vendor Heterogeneity: is any per-vendor outcome left once PR size is matched, given Claude Code's 495-line median?
- Verification as the New Bottleneck: how far do you push fully automated reviews?
- The Three Loops of AI-Native Building: is Ng's "QA burden fell significantly" a 0-to-1-vs-production scope split, or optimism bias?
One of the six retires. The other five each come down to a number the wiki does not hold, so they get partial answers and move to #oq/source.
Answer 1: acceptance of an agent-applied change is survival, and it needs two qualifiers#
Faros's 20%→60% (65% in the successor report) is "the overall rate of AI-generated code being accepted into codebases," driven "substantially by tools like Cursor and Claude Code operating in agent mode, where the agent applies changes directly" (AI Engineering Report 2026: The Acceleration Whiplash, §acceptance). Faros never says what counts as acceptance when nobody performs an accept action. The corpus now lets the construct be stated completely.
For an agent-applied diff, "acceptance" means non-reversion, which is a survival measure rather than an act. It needs two qualifiers before it says anything about oversight:
- A horizon. DECODE shows that even affirmative acceptance keeps changing: 31% of trajectories contain a removal edit, and the median completion loses about a third of itself within the hour. Survival without a stated window is undefined. Kraishan shows what a defined window looks like: reverts within 90 days of merge. The same page names what a revert-commit detector misses. It misses silent rewrites, so a change undone in substance but not in the commit message counts as surviving. Non-reversion measured that way overstates acceptance.
- An examiner class. The earlier synthesis Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? split the construct into affirmative adoption, reviewed non-reversion and bare non-reversion. Closed-Loop AI Review adds a fourth class the partition lacked: AI-reviewed non-reversion, where the agent applies and another agent looks. That is 248,641 of 2,830,284 attributable agent-authored PRs (8.8%), and 83.7% of those were reviewed by the authoring product itself. The same page warns that a GitHub-mined "AI-reviewed" record does not exclude a human reviewer, and "no AI reviewer" does not exclude AI review. So the four classes are well defined but not cleanly observable from mined PR data. Review as the Control Point shows the invoker class is large: author-only review covers 40.1% of agent PRs against 21.5% of human PRs.
The best instrument for the construct is authoring-time provenance. Tran et al. measure byte-level provenance at submission, and the paper notes that developers "substantially filter generated text before it reaches submitted-code analysis." That is survival of the human filter, measured without any accept event, which is exactly what acceptance means once the agent applies the change.
So the question as posed has an answer: acceptance of an agent-applied change = the diff's survival over a stated horizon, stratified by who examined it (nobody / the invoker / an independent human / an AI reviewer), and, where revert detection is by commit message, an upper bound on survival. The 60% is uninterpretable as oversight because it publishes none of the three parameters. What stays open is a different question: the shares of the four classes on any single population. The pieces exist only on incommensurable corpora: 40.1% invoker-only on Agents in the Wild, 8.8% AI-reviewed on CodAGE, Faros's 31.3% merged-unreviewed growth on its customer panel. That shares question is filed as a new #oq/source question on AI as Primary Author.
Answer 2: the quarantined-PR review rate is not in the wiki, and the hole is larger than the quarantine#
The paper reports no AI-review rate for the 1,733,535 quarantined author-side PRs. It reports only the quarantine size and reasons: 38.0% of candidates, 96.9% branch_only (AI-to-AI Code Reviews of GitHub Pull Requests §3.3). The number cannot be synthesized. It needs a hand-labelled sample drawn from the released quarantine files, which makes this a new-evidence question.
The synthesis adds one point the question's framing misses. The quarantine is not the whole denominator hole. The construct-validity section says unsigned PRs "are not quarantined but simply never enter our population" (§7). A PR with no body signature, no vendor login and no branch prefix is invisible to the pipeline. Labelling a sample of the 1.73M quarantined PRs would bound the review rate on the branch-prefixed part of the missing population, not on all AI-authored PRs. An ecosystem rate would need a second sample drawn from outside the signature framework altogether.
Answer 3: the comment gap is already reviewer-fixed, and holding the reviewer fixed moves the confound to the author#
The question assumes the 58–65% gap is confounded with reviewer identity. That holds for the pooled same/cross contrast and for the latency result, where the page shows the reviewer breakdown explains 99.8% of the aggregate. It does not hold for the comment-volume figures. Each row of that table is one reviewer bot comparing its own product's PRs with other products' PRs: Copilot 1.49→2.35, Devin 1.09→1.80, Amazon Q 4.94→8.08, Codex 0.91→0.89 (AI-to-AI Code Reviews of GitHub Pull Requests line 239; Closed-Loop AI Review §"What moves with the pairing"). So the reviewer-fixed half of the question is answered by construction: with the reviewer held fixed, the gap is present for three of four bots and absent for Codex. The Codex row is therefore not a hint that the pooled gap is confounded. It is the one reviewer whose own gap is flat.
Holding the reviewer fixed does not remove the confound. It moves it. With the reviewer fixed, the thing that varies is the authoring product. Of Copilot's 21,022 cross-product PRs, 18,114 (86%) are Codex-authored, the modal pair in the whole dataset. So Copilot's 58% gap is close to "Copilot reviewing Copilot-authored PRs vs Copilot reviewing Codex-authored PRs." That contrast carries every difference in size, task and repository between the two authors' PRs. The paper lists change size among the characteristics that differ between the arms and declines any causal reading (§5, §7 Internal validity). It does not match on size, and its own future-work design (matched PRs "aligned on repository, language, task type, and change size") is exactly the missing control.
No size information on CodAGE PRs exists in the wiki. The AIDev medians on Agent-Vendor Heterogeneity cannot stand in for it: the two corpora are explicitly marked not comparable. The size half is unanswerable from the wiki. It needs the released data re-analysed, which is new evidence.
Answer 4: the headline vendor gap survives a median-size comparison; the Claude Code anomalies are the size-sensitive ones#
The paper runs no size-matched per-vendor analysis. Its only size control is a size-bucket check on the pooled agent-vs-human security density (difference confined to the XL bucket, δ = −.08) (Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild RQ1; §5 "Size is the confound to watch"). What the wiki does hold, Table 1's median sizes against Tables 2–4, is enough to split the page's results into two groups:
| Result | Groups compared | Median changed lines | Reads as size artifact? |
|---|---|---|---|
| Revert OR 0.50 vs 1.31 (CIs [0.44, 0.57] vs [1.11, 1.54]) | Codex vs Devin | 63 vs 61 | No: near-identical medians, non-overlapping CIs |
| Comment ratio.065 vs 0 at the median | Copilot vs Codex | 76 vs 63 | Unlikely: a 13-line gap against a presence/absence difference |
| Maintainability-smell density δ = −.19 vs human | Codex vs human | 63 vs 52 | No: already per-line |
| ≥1 security smell 9.5% vs 4.6% | Claude Code vs human | 495 vs 52 | Yes, plausibly: per-line density vs humans is negligible (δ = +.05) |
| Heaviest structure (nesting 4 vs 3) | Claude Code vs rest | 495 vs 52–96 | The paper itself calls it "consistent with its much larger PRs" |
| First human review 12.6 h vs 1–4 h | Claude Code vs rest | 495 vs 52–96 | The paper offers size as the explanation |
So the question's fear that a null "would collapse most of this page into a statement about PR size" is too broad. The page's headline result, Codex vs Devin on reverts, compares two groups whose median PRs differ by two lines. Claude Code's three anomalies are the ones a size explanation fits, and its presence-rate outlier sits beside a negligible per-line density.
This is partial, not settled, for two reasons. Median matching is not distribution matching: a fat Devin tail could still carry its revert excess, and revert odds plausibly rise with size. And the revert detector reads the most-changed file only. A size-matched resampling on the released pipeline is still the test, and the wiki does not contain one.
Answer 5: "how far" gains a "by whom" axis, and the default is the disfavoured configuration#
The existing annotations already answer "how far" as a partition: automate mechanical checking fully, contest the quality/security edge (P9), and protect the non-defect functions of review. Two things in this cluster add to it.
The ecosystem default for automated review is self-review. On public GitHub, same-product AI review outnumbers cross-product 208,145 to 45,269, about 4.6:1 (Closed-Loop AI Review). The only recall measurement the wiki holds for automated reviewers finds each model catching 6–12 points fewer high-severity bugs in its own family's code (Same-Model Review Blindness, case-study). The two measure different constructs (product pairing vs model lineage, comment volume vs recall), so this is a tension to record, not a finding. But it means "how far" has a routing dimension the partition lacked: how much review is automated, and also whether the automated reviewer is independent of the author. Where it is automated today, it mostly is not. Meanwhile 91% of attributable agent PRs (2,581,643 of 2,830,284) drew no detectable AI review at all.
A size-keyed partition line is also a vendor filter. PostHog's <500-line/<20-file ceiling (Risk-Tiered Auto-Approval) passes the median Codex, Devin, Copilot and Cursor PR and blocks essentially every Claude Code PR, whose median is 495 lines (Agent-Vendor Heterogeneity). Where the line is drawn by cheap structural checks, it silently routes by authoring tool. By Answer 4, the tool gaps that survive a median-size comparison (Codex vs Devin reverts) are ones a size gate cannot see.
The safety half is still missing: a false-approval or escaped-defect rate for automated review at production scale. ICONIQ's 0 of ~180 autonomous-agent PRs passing CI and a security leader's "20 to 30% false-positive rates [are] a bigger tax on the team than finding more issues" (2026 State of Scaling: The Great Sorting p.46, practitioner-opinion) are the cost-side data points. Neither is an escape rate. Only a new source can supply one.
Answer 6: Ng's own mechanism is the weak link#
Ng's claim is specific about its cause: "with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly" (Thread by @AndrewYNg). The earlier synthesis answered both-and: the scope split is real, and self-report can't be read as magnitude. This cluster adds evidence against the stated mechanism, which is a third reading besides "scope" and "optimism."
Dipongkor et al. (empirical, 4,882 agentic PRs) find that agent-written tests raise coverage of the agent's own diff in only 35.9% (Java) / 22.5% (Python) of the PRs that include tests. 50.4% of code-changing PRs carry no test change, and error-handling constructs go unexercised up to 86.0% of the time. Verification as the New Bottleneck draws the consequence: a green suite that never touched the diff is no evidence dressed as the strongest kind. So a felt fall in QA time is consistent with less checking, not only with less need for it. The human stopped manually testing, and the agent's self-tests often do not point at the change.
This reading is weak for Ng's actual setting. Dipongkor measures open-source PRs against existing suites. In a 0-to-1 build, every line is new and the "existing suite" is whatever the agent wrote, so the diff-coverage gap may be smaller. The question still needs what the earlier answer named: a measured QA-time (or escaped-defect) series for 0-to-1 builders. No wiki source has one, so it is a new-evidence question.
Verdicts#
| # | Page | Verdict | Why |
|---|---|---|---|
| 1 | AI as Primary Author | Retired | The construct question is answered: survival over a horizon × examiner class (four classes), with revert-by-message as an upper bound. The measurement residual (class shares) is split into a new #oq/source question |
| 2 | Closed-Loop AI Review (quarantine rate) | Partial → #oq/source | The number is absent from the paper and the wiki. It needs a labelled sample, and the quarantine undercounts the hole |
| 3 | Closed-Loop AI Review (same-product gap) | Partial → #oq/source | The reviewer-fixed half holds by construction (3 of 4 bots), but the author product is then confounded. The size half needs re-analysis |
| 4 | Agent-Vendor Heterogeneity | Partial → #oq/source | The Codex–Devin revert gap sits on matched medians (63 vs 61). No size-matched resampling exists |
| 5 | Verification as the New Bottleneck | Partial → #oq/source | Adds the by-whom axis and the size-gate-as-vendor-filter point. The safety half needs an escape rate |
| 6 | The Three Loops of AI-Native Building | Partial → #oq/source | Ng's mechanism is undercut by diff-coverage data, but no 0-to-1 QA-time measurement exists |
Cited by 8
- AI as Primary Author×3
Agent Pr Oversight Numbers And Their Confounds — retires the "acceptance" question: survival over a…
- Closed-Loop AI Review×3
Same-product and cross-product are nearly confounded with reviewer identity: the same-product arm…
- Open Questions Backlog×3
Closed Loop Ai Review: Same-product and cross-product are nearly confounded with reviewer identity:…
- Agent-Vendor Heterogeneity×2
Agent Pr Oversight Numbers And Their Confounds — sorts this page's results by size exposure: the…
- The Three Loops of AI-Native Building×2
Agent Pr Oversight Numbers And Their Confounds — Ng's QA-relief claim tested against its own stated…
- Verification as the New Bottleneck×2
Fung's own open question: "How far do you push fully automated reviews?" — where's the speed/safety…
- Follow-Up Fixes on Agent PRs
Agent Pr Oversight Numbers And Their Confounds — that synthesis treats survival with…
- AI Coding Practice
Agent Pr Oversight Numbers And Their Confounds — Six-question synthesis of the agent-PR oversight…
Related articles
- Review as the Control Point
Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Security Debt of Agent-Generated Code
Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…
- Agent-Generated Test Quality
Two AIDev cuts on whether agent code is tested. Jhanglani et al. (204K test files): a trade, not a deficit — agents dou…
