Sources#
Summary#
Every other review measurement in this vault is an experiment, a vendor's labelled set, or a single corpus of a few thousand PRs. Selvanayagam & Ghaleb (École de technologie supérieure Montréal and Trent University, arXiv 2608.21311, ESEM 2026 Emerging Results, Vision & Reflection track) do the census instead: how often, across all of public GitHub, does one identifiable AI coding product open a pull request that another identifiable AI coding product reviews?
The answer is that the configuration is already ordinary and still a minority. Of 2,830,284 signature-attributed agent-authored PRs from 2024-01-01 to 2026-04-15, 248,641 (8.8%) received at least one AI-attributed review. Within those: 45,269 cross-product, 208,145 same-product, 4,773 both — and the three close exactly (45,269 + 208,145 − 4,773 = 248,641), which is the cleanest internal check the paper offers and it does not state it. Cross-product review is 1.6% of agent-authored PRs and grew by more than two orders of magnitude from 2025-Q1 to 2025-Q3 (57 → 25,492 PRs).
What the paper is not is a quality measurement. It measures comment counts, self-declared comment categories and timestamps; it has no defect ground truth, no human-authored control cohort, and no outcome linkage. Its own closing line: "These findings describe observable reviewer behavior, not review correctness or software quality." Read it as the denominator and the composition that the vault's efficacy measurements — Same-Model Review Blindness, Agent Review Comment Resolution, Deterministic Engineering for Agent Code Review — have been quoting rates against without knowing the population.
Evidence note.
empirical, tier kept. Secondary analysis of a public third-party dataset (CodAGE, CC BY 4.0 on Hugging Face) with a released replication package: signature registry, classifier rule set and ordering, processing scripts. No vendor affiliation and no product. The limits are the ones a mining study has — everything is observational, nothing is controlled, and it is a short emerging-results paper, so the analysis stops where a full paper would start validating. The authors are unusually disciplined about saying so: the counts are declared lower bounds in three separate places, the latency comparison is explicitly withdrawn from causal reading, and §5.1 refuses to interpret CodeRabbit's own labels as a severity scale even though the released artifact calls them one.
Who is actually in the population (the part that bounds everything else)#
Attribution is a two-tier signature framework, and the tiers are the reason the author-side and review-side numbers are not equally trustworthy.
- S1, body signature — a string the agent itself emits: the
Co-Authored-By: Claude <noreply@anthropic.com>trailer from Claude Code, thecursor.com/agentsURL from Cursor, thechatgpt.com/codexURL from OpenAI Codex. S1 agents require body evidence. - S2, vendor-controlled login —
coderabbitai[bot],devin-ai-integration[bot],gemini-code-assist[bot]. S2 agents accept body or login evidence. - Branch prefixes are recorded and never sufficient (
codex/,cursor/), being user-controllable.
The consequence is an asymmetric coverage hole, and it is enormous on exactly the side the headline rate divides by. Quarantine is 0.01% on both review streams (590 of 4,141,598 review events; 793 of 8,560,919 review comments) — the major reviewers post under vendor-controlled bot logins, so the reviewer side is nearly complete. On the author side, 1,733,535 of 4,563,819 candidate PRs (38.0%) are quarantined, 96.9% of them as branch_only. The agent-authored population is the 62.0% that carried a body or login signature.
So the "8.8% of agent-authored PRs get AI review" and "1.6% get cross-product review" figures have a near-complete numerator over a demonstrably incomplete denominator, and the paper says which way that cuts: absolute counts are lower bounds, rates "inherit the coverage of both sides and should be read as describing the attributable population rather than all AI-authored PRs." If the 1.7M quarantined branch-only PRs are AI-authored at any appreciable rate and are no less likely to draw AI review, the true rate is lower than 8.8%; if signature-emitting agents are also the integration-heavy ones, it is lower still. Nothing here supports treating 8.8% as an ecosystem-wide review rate.
Two more filters: agents with fewer than 50 attributed PRs (Aider, Kiro, SWE-agent, Windsurf) are dropped from per-agent claims, and 98 review entries and 111 review-comment entries (0.002% / 0.001%) matched two agents at once — one agent quoting another's trailer, which is the attribution hazard this design was built against showing up at a negligible rate.
Prevalence, growth, and the composition flip#
Growth is the headline and it is a 2025 phenomenon. Activity was negligible through 2025-Q1 (57 cross, 40 same) and reached 25,492 cross and 57,080 same by 2025-Q3. 2025-Q4 is hatched in Figure 1 as a lower bound because GHArchive attribute lag runs about a quarter. The closed-loop population spans 10,345 repositories — the RQ1 summary's antecedent is ambiguous between the 248,641 and the 45,269, and the paper never disambiguates it, so read it as the order of magnitude rather than a per-arm figure.
Two readings off Figure 1 that the prose does not make. First, the two curves cross: in 2025-Q2 cross-product review ran ahead of same-product (roughly 10⁴ against 4×10³ on the log axis) and same-product overtook it in Q3 by better than 2:1. Whatever drove the Q3 step is therefore a same-product event, and given that Copilot is 80% of the same-product arm it is most likely something about Copilot's own review surface rather than about closed-loop review in general — the same integration effect Kraishan invokes to explain Copilot PRs drawing the most bot reviews of any vendor. The paper does not decompose the growth by product, so this is a reading, not a finding. Second, 2024 is not empty but it is noise — tens of PRs per quarter, with 2024-Q3 showing a cross-product bar and no same-product bar at all.
Note what the paper's own framing does to the arithmetic: it focuses on cross-product review "because same-product review may be part of an integrated product workflow," which sets aside 83.7% of the closed-loop population (208,145 of 248,641). The interesting fraction is a fraction of a fraction, and the boring majority is the one growing fastest.
The pairing is not uniform — it is three separate regimes#
Table 2 (reconciled against pdftotext -layout; the cross and same columns sum to their stated totals in all eight rows, and the same-product column sums to exactly 208,145) splits authoring agents into three groups that share almost nothing:
| Authoring agent | Cross-product PRs | Same-product PRs | Total A×B | % same-product |
|---|---|---|---|---|
| Copilot | 7,515 | 166,442 | 173,957 | 95.7% |
| OpenAI Codex | 31,601 | 37,247 | 68,848 | 54.1% |
| Devin | 1,171 | 3,811 | 4,982 | 76.5% |
| Cursor | 2,309 | 0 | 2,309 | 0.0% |
| Claude Code | 2,042 | 10 | 2,052 | 0.5% |
| Amazon Q | 48 | 535 | 583 | 91.8% |
| Google Jules | 478 | 0 | 478 | 0.0% |
| Sweep AI | 100 | 100 | 200 | 50.0% |
- Product-internal (Copilot 95.7%, Amazon Q 91.8%, Devin 76.5%): the vendor ships both halves and the loop closes inside one product.
- Balanced (Codex, 54.1%): the only author with real population on both sides.
- Structurally cross-product (Cursor, Google Jules, Claude Code): these products have no reviewer, so every review their PRs receive necessarily comes from someone else. Cursor's 0.0% and Jules's 0.0% are not a preference; they are a product-surface fact.
That last group matters for how the same/cross contrast can be read at all. The cross-product arm is heavily populated by agents that could not have self-reviewed, and the same-product arm is 80% Copilot. "Same vs cross" is therefore very close to "Copilot vs everyone else" with a Codex column in the middle — a confound the paper names as reviewer composition but never quantifies. On the reviewer side the concentration is the same: Copilot is 21,022 of the 47,259 cross-product author-reviewer pairs, and the single modal pair in the whole dataset is Codex authored, Copilot reviewed — 18,114 pairs, 38% of all cross-product pairs on its own.
One detail worth keeping for anyone reusing the corpus: product boundaries are drawn at the agentic harness, not the company. Google Jules and Gemini Code Assist count as different products, and 117 of the 45,269 cross-product PRs (0.26%) are that kind of same-vendor crossing. The categories deliberately do not track corporate vendor, and they explicitly do not track foundation model.
What moves with the pairing, and what it is evidence of#
Three measures, all reviewer output rather than review effect.
1. Comment-category mix moves a lot, with the reviewer held constant. Holding CodeRabbit fixed and varying the author across 35,248 classified comments, the refactor share runs from 9.7% (Devin-authored) to 35.0% (Claude Code-authored) and potential_issue inversely, 29.9% to 49.0%. The largest contrast, and the one the paper defends: 35.0% refactor on Claude Code PRs against 10.5% on Copilot PRs, a 24.5pp difference, 95% CI [23.1, 25.9]. The association is significant and small — Cramér's V = 0.150, χ² = 3177.5, df = 24. The authors' own reading is the correct one: this is reviewer output composition, and the PRs themselves (change size, language, repository) are the obvious alternative explanation. It is not a code-quality claim, and the labels are CodeRabbit's self-declared headers, deterministically extracted, never validated against anything.
2. Comment volume is higher on same-product PRs — the direction the blindness framing would not predict. For the four bots with ≥200 PRs on both sides:
| Reviewer bot | Cross PRs | Mean cross | Self PRs | Mean self | p | |δ| |
|---|---|---|---|---|---|---|
| Copilot | 21,022 | 1.49 | 166,442 | 2.35 | < 10⁻²⁹² | 0.15 |
| OpenAI Codex | 5,485 | 0.91 | 37,247 | 0.89 | 2.0 × 10⁻³ | 0.02 |
| Devin | 392 | 1.09 | 3,811 | 1.80 | 3.0 × 10⁻⁶ | 0.14 |
| Amazon Q | 522 | 4.94 | 535 | 8.08 | 2.2 × 10⁻¹⁴ | 0.27 |
Three of four bots write 58–65% more comments per PR on their own product's code; Codex is flat. The effect sizes are the story: Cliff's δ is negligible-to-small throughout, Devin and Copilot sit within 0.01 of the 0.147 boundary, and the authors state plainly that the mean gap is "not a population-wide shift but a long right tail of high-comment same-product PRs." A p < 10⁻²⁹² on n = 187,464 is a sample-size artifact and they say so.
3. Latency differences are almost entirely disqualified by their own missingness. Median time from PR creation to first AI review: 1.2 minutes cross-product, 4.7 minutes same-product — counter-intuitive, and the paper immediately dismantles it. Timestamp retention is 79.2% of cross-product pairs against 31.9% of same-product pairs, which is not a nuisance rate but a different sample. And the reviewer-level rows explain the aggregate without any product effect: Gemini Code Assist's median is 0.5 min and Copilot's is 17.7 min (p75 73.9 min), and Copilot dominates the same-product arm. The reviewer rows sum to 103,743 of the 103,920 qualifying pairs (99.8%), so the reviewer breakdown is effectively exhaustive and the compositional explanation is complete — there is no residual for product pairing to explain. The one durable finding here is the uncontroversial one: AI review typically begins within minutes, against no human baseline.
The construct gap: product is not model#
This is the load-bearing caveat for every use the vault will make of this page, and the authors raise it themselves three times — in §3.4, in External Validity ("'cross-product' reflects product-level attribution rather than model-level identity"), and as the first item of future work.
Public GitHub records reveal the product that opened or reviewed a PR and essentially never the model behind it. So:
- A same-product pair need not be a same-model pair. A product may route to several models, or change models between the authoring and reviewing surfaces, or between releases.
- A cross-product pair need not be a cross-model pair. Two products can sit on the same foundation model — the paper names this directly and the vault's own Agent-Vendor Heterogeneity page is built on vendor labels with the same blind spot.
- The 117 same-vendor cross-product PRs make the point in miniature: even the corporate boundary is not the product boundary, let alone the model boundary.
The consequence for Same-Model Review Blindness is precise and limited. That page's finding is a lineage effect measured on ~1,500 labelled high-severity bugs — a reviewer catches 6–12 points fewer of them in its own family's code. This paper measures a harness-pairing effect on comment counts in a population 100× larger with no ground truth at all. The comment-volume result runs the other way (more comments when the product reviews itself, not fewer findings), but it is not the same quantity, not the same construct, and not a contradiction: volume is not recall, and the same-product arm is dominated by one product whose reviewer sits inside GitHub's own review surface. What this source genuinely supplies is the population the routing prescription would operate on — 208,145 PRs a year where a product reviewed its own output, against 45,269 where it did not — and the observation that the ecosystem's default is currently self-review, by better than four to one (208,145 against 45,269).
Why this is a caveat on the vault's own evidence base#
The paper's sharpest implication is methodological and points straight at the corpus this wiki is built from: "studies must explicitly account for AI-generated artifacts or risk mixing fundamentally different populations."
Several pages here rest on mined GitHub review data — Agent Review Comment Resolution (54,713 agent review comments in 341 repositories), Agent-Vendor Heterogeneity (review behaviour across 33,596 vendor-labelled PRs), Review as the Control Point (2.5M+ PRs of non-vendor telemetry). Each treats the review stream as, at minimum, interpretable. This source says that on agent-authored PRs the review stream contains hundreds of thousands of AI-attributed events by 2025-Q3, and — critically — its own "closed loop" does not mean humans were absent: the review stream mined is restricted to AI-attributed events, so a closed-loop PR may also have had a human reviewer nobody can see. The contamination runs both directions. A GitHub-mined "review" is not evidence of human decision-making, and a GitHub-mined "no AI reviewer" is not evidence that no AI reviewed.
That is a bound on instrument, not a refutation of any particular finding, and it applies most sharply to any measurement that reads review presence, review latency or review volume as a proxy for human attention.
Connections#
- Same-Model Review Blindness — the population this page's census describes, and the construct that census cannot reach. That page measures a reviewer losing 6–12 points of high-severity recall on its own model family's code, on ~1,500 vendor-labelled bugs; this counts 208,145 PRs where a product reviewed its own output against 45,269 where it did not, with no ground truth and product-level rather than model-level attribution. The one place their numbers touch runs the other way — three of four dual-role bots emit 58–65% more comments per PR on same-product code — but comment volume is not defect recall, the same-product arm is 80% one product, and Cliff's δ is negligible-to-small. What transfers cleanly is scale and default: the ecosystem's modal closed-loop configuration is currently self-review, by better than four to one, so the routing prescription there is a minority practice, not the status quo
- Agent-Vendor Heterogeneity — the same "the vendor label is the variable" instruction, arriving from the reviewing side at 75× the corpus size. Kraishan's spread is in outcomes (revert odds 0.50 to 1.31 against one human baseline) across 37,623 AIDev PRs; this one's is in the review configuration itself (95.7% self-reviewed to 0.0%, across 248,641 CodAGE PRs) and in reviewer output (CodeRabbit's refactor share 9.7% to 35.0% by author). Both inherit the same blind spot and both declare it: the label names a product, and the product is a bundle of model, harness, integration surface and the tasks its users hand it. The two corpora are distinct and not comparable — AIDev is 33,596 curated PRs from 2,807 >100-star repos, CodAGE is a GHArchive-wide event stream — so no figure from one may be read against the other
- Review as the Control Point — the second moderator's denominator, and a measurement hazard for the theory's own telemetry half. The theory treats automated-reviewer capability as a construct without saying how much of the population it covers; this says 8.8% of attributable agent-authored PRs by 2026, growing two orders of magnitude in three quarters. It also sharpens the warning under that page's 2.5M-PR observational study: review events on agent PRs are increasingly AI-attributed, and AI-attributed review does not exclude human review, so neither "reviewed" nor "unreviewed" is a clean read on human attention any more
- Agent Review Comment Resolution — the same loop, one construct apart, and they compose. That page measures what humans do with 54,713 agent review comments (71.4% resolved); this measures how many such comments exist and who they come from, at a scale where the reviewer streams alone hold 8,560,237 attributed review-comment entries. Neither touches correctness. Read together they say the layer is large, growing, adopted at roughly seven in ten — and entirely unmeasured on whether any of it prevents a defect
- Deterministic Engineering for Agent Code Review — the benchmark-side complement. OpenCodeReview scores review systems against 1,505 expert-verified comments on a curated benchmark; this counts what those systems actually emit in the wild and finds the output mix moves 24.5pp with who wrote the code. The pairing is a variable no review benchmark in this corpus holds fixed or reports
- Open Source Under Agent Contributions — the maintainer-facing number behind DHH's account of unbounded contribution supply. 2.83M attributable agent-authored PRs, of which 2,581,643 drew no detectable AI review at all — so the first-pass-triage-by-agent practice he describes is, at population scale, not what most agent PRs get
- Telemetry vs. Survey Measurement — the mined-artifact version of that page's definition problem. There, an adoption rate moves 40× with the threshold; here, an "agent-authored PR" count moves by 1.73M depending on whether a branch-name prefix counts as evidence, and the paper's choice to exclude it (96.9% of all quarantine) is the single largest methodological decision in the study
- Verification as the New Bottleneck — the bottleneck's automated half, sized. Whatever share of verification is being delegated to an AI reviewer, the measurable floor on public GitHub is 248,641 PRs and rising two orders of magnitude a year
Open Questions#
- The 8.8% review rate divides a near-complete numerator (0.01% reviewer-side quarantine) by a demonstrably incomplete denominator (38.0% author-side quarantine, 96.9% of it branch-only). What is the AI-review rate on the 1.73M quarantined branch-only PRs? A hand-labelled sample of a few hundred would bound it, and the replication package plus the quarantine files with recorded reasons make it a day's work. Until then no rate on this page may be quoted as an ecosystem rate rather than a rate over the signature-attributable population.
- Same-product and cross-product are nearly confounded with reviewer identity: the same-product arm is 80% Copilot, the cross-product arm is heavily populated by Cursor, Google Jules and Claude Code, which have no reviewer of their own and so could not have self-reviewed. Does the 58–65% same-product comment-volume gap survive holding the reviewer bot fixed and matching on change size? The Codex row already hints not (0.91 vs 0.89 on 42,732 PRs, the one bot with substantial population on both sides), and the data to test it is released.
- The paper's own first future-work item is the experiment this vault most wants: review functionally equivalent PRs from known models under fixed configurations with source blinding, scoring defect recall and false positives across same-model, related-model and independent-model reviewers. Does the product-pairing effect measured here survive when the model behind each product is known and held fixed? Nothing observable on public GitHub can answer it — the records do not carry model identity — so this is a controlled-experiment question, not a mining one.
Sources#
- AI-to-AI Code Reviews of GitHub Pull Requests — Niruthiha Selvanayagam (École de technologie supérieure, ÉTS Montréal) & Taher A. Ghaleb (Trent University), AI-to-AI Code Reviews of GitHub Pull Requests, arXiv 2608.21311, 2026-08-21, 15 pages, DOI 10.4230/LIPIcs.ESEM.2026.74, ESEM 2026 Emerging Results, Vision & Reflection track.
empirical, tier kept. Ghaleb is also the author of CodAGE [7] and of the agent-fingerprinting work [8] this paper's attribution framework extends, so the dataset and the study share an author — disclosed here because it is the nearest thing to a conflict on the page, and it cuts toward care rather than away from it (the quarantine and coverage reporting are unusually explicit for a short paper). Sections used: §3.1–3.3 (CodAGE snapshot 2024-01-01 to 2026-04-15, the S1/S2 signature tiers, quarantine rates and the branch-only asymmetry), §3.4 (the three datasets and the product-not-vendor-not-model definition, including the 117 same-vendor cross-product PRs), §3.5 (CodeRabbit's self-declared headers and the ordered-regex classifier), §4.1–4.3 + Figure 1 + Tables 1–2 (prevalence, growth, crosstab, composition), §5.1–5.3 + Figure 2 + Tables 3–5 (category mix, comment volume, latency), §6 (the sampling-contamination implication), §7 (all three validity dimensions), §8 (the three future-work programmes). Parse status: clean, verified in full. PDF-derived (docling 2.126.0,docling_mlx0.1.1, MLX layout and table stages, rapidocr,confidence_grade: excellent). Ingestverifyreturnedwarnon atable-collapseflag against Table 4'spcolumn; the flag is a confirmed false positive — docling drops the superscript minus on scientific notation, so10⁻⁵-shaped values render as10 - 5and trip the multi-value-cell heuristic. A> [!note]block in the raw records this. All five tables were then reconciled cell-for-cell againstpdftotext -layouton pages 7–10 and match exactly — no collapse, no shift, no weld, no split row, no dropped rows, no en-dash corruption, noAl-for-AI. Four arithmetic checks were run on top and all close: Table 2's cross and same columns sum to their stated row totals in all eight rows and its same-product column sums to exactly 208,145; Table 1's row totals reproduce (32,379 / 7,355 / 2,109 / 2,173 / 656); Table 3's category shares sum to 100.0% per row and its N column to the stated 35,248; Table 4'sSelf PRscolumn matches Table 2's same-product totals for all four bots and the stated 58/65/64% gaps reproduce from the means. Table 5's reviewer rows sum to 103,743 of the stated 103,920 qualifying pairs (99.8%, the remainder being bots below the ≥100 threshold). Images: all five assets viewed under the two-pass rule —image_000002is Figure 1 (quarterly log-scale bars, source of the 2025-Q2 crossover read on this page),image_000003is Figure 2 (the stacked category bars, which independently reproduce every Table 3 value);image_000000andimage_000001are the CC-BY and LIPIcs logos andimage_000004is a section-number badge — all three decorative, none load-bearing. Two source-internal oddities, neither a parse artefact. (1) The Table 5 caption is doubled in the PDF itself — a bareTable 5line immediately followed by the real caption labelledTable 6, in a paper with five tables; a LaTeX slip, and the table's content and numbering are otherwise consistent. (2) The snapshot is stated twice as running to 2026-04-15, yet Figure 1's axis stops at 2025-Q4 and marks that quarter as a lower bound for "about one quarter" of GHArchive attribute lag — a window that should have closed by April 2026. The paper never reconciles the two; nothing on this page is cited from 2025-Q4 or later.
Cited by 10
- Review as the Control Point×3
The theory's second moderator — automated-reviewer capability — has never had a denominator on this…
- Same-Model Review Blindness×3
ai to ai code reviews — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21,…
- Agent Review Comment Resolution×2
ai to ai code reviews — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21,…
- Agent-Vendor Heterogeneity×2
ai to ai code reviews — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21,…
- Deterministic Engineering for Agent Code Review×2
ai to ai code reviews — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21,…
- Open Questions Dashboard×2
Closed Loop Ai Review: Same-product and cross-product are nearly confounded with reviewer identity:…
- Open Source Under Agent Contributions×2
ai to ai code reviews — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21,…
- Telemetry vs. Survey Measurement×2
Closed Loop Ai Review — the definition problem on mined artifacts rather than on survey items. This…
- AI Coding Practice
Closed Loop Ai Review — Selvanayagam & Ghaleb (ÉTS Montréal / Trent, arXiv 2608.21311, ESEM 2026…
- Open Questions Backlog
Closed Loop Ai Review ×3 (oldest 7d) — The 8.8% review rate divides a near-complete numerator…
Related articles
- Security Debt of Agent-Generated Code
Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Agent Review Comment Resolution
Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex…
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Efficiency Debt of AI-Generated Code
Tran et al. (Google, arXiv 2608.06640): 3.52M changes over 12 months in one production C++ monorepo, with a human-writt…
