H
Howardism
Plate IIAI Coding PracticeHOWARDISM

Follow-Up Fixes on Agent PRs

Takerngsaksiri, Duong & Barnett (Deakin, arXiv 2609.26847): merged agent PRs draw a verified follow-up fix within 30 days at 1.62x the odds of same-repo human merges (3.68% vs 2.34%), 69.6% of those fixes come from the same agent and 89.1% of fixes to human merges from humans. This is the first controlled post-merge measure in the vault to put agents above human parity, because it counts forward fixes, the channel revert-by-message studies miss. Most of the excess arrives on the day of the merge, and 'self-fix' means same product, not same session

Article metadata
Publication details
Published:October 1, 2026
Filed:Concept
Domain:AI Coding Practice
Reading:17 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Follow-Up Fixes on Agent PRs

Sources#

Summary#

A merged pull request usually counts as finished work. Takerngsaksiri, Duong & Barnett (Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests, Deakin A2I2, arXiv 2609.26847, v1 2026-09-22, empirical) follow merged agent PRs past the merge and ask two things: how often a later PR has to fix one, and who writes that fix. The corpus is AIDev-pop (repositories with ≥500 stars): 6,774 merged agent PRs across 891 repositories from five agents (OpenAI Codex 2,834, Devin 1,813, GitHub Copilot 1,428, Cursor 569, Claude Code 130), set against 5,044 human PRs from the same repositories.

The headline results:

  1. Agent merges need fixing more often. In the shared observation window, 3.68% of agent merges draw a verified fix within 30 days against 2.34% of human merges. The Mantel-Haenszel odds ratio, stratified over the 218 repositories where both cohorts appear, is 1.62 [1.10–2.39], p = 0.015, and the risk ratio is 1.59.
  2. Whoever wrote the code fixes it. 69.6% of verified fixes on agent merges are opened by the same agent, 27.4% by a human and 3.0% by another agent (n = 263 pairs). On human merges, 89.1% of fixes are human (n = 110).
  3. The fixes are agent-written, not just agent-opened. On agent-opened fix PRs (Codex excluded, n = 123), commits are on average 87.4% agent-authored per PR, and 76.4% of those PRs are agent-authored in every commit. Across all merged agent PRs the same figures are 78.8% and 54.1%.

The authors conclude that "agents currently largely finish their own job, but their merges still require fixing more often than human merges."

How a fix is identified#

The construct matters more than usual here, because it is the reason this paper disagrees with the vault's revert studies (see below).

  • Candidate fix. A later PR Q counts as a candidate fix for a merged PR P when all four filters hold. Q merges in the same repository within 30 days after P. AIDev's task-type tag labels Q a fix. Q's net diff co-edits at least one file P merged. At least one of those shared files is not boilerplate (lockfiles, changelogs and generated artifacts are excluded). Merges less than 30 days before the dataset end are dropped as partially observed. The same filters run on the human cohort.
  • Verified fix. A candidate pair is verified only if it is labelled Direct fix: Q repairs a defect or incomplete behaviour that P introduced, and modifies or reverts the lines P added. The alternatives are Related touch, Unrelated or Indeterminate. Two annotators agreed on 45 of 50 pairs (binary κ = 0.77). An LLM judge (Claude Opus 4.8, temperature 0), built on the 45 agreed pairs, labelled the full population of 2,510 agent-merge and 2,206 human-merge pairs. It re-validated at κ = 0.78 against a fresh human relabel, with Direct-fix precision of 27/30 on both cohorts (pooled 54/60 = 90%, Wilson 80–95%).
  • The label is a floor. By the paper's own account the linker misses cross-file fixes, fixes slower than 30 days, and fixes folded into PRs not typed as fixes. Because both cohorts share the linker, the comparison is fair even though both rates are low.

The judge validation is careful by the standard LLM-Judge Validation sets. There is a chance-corrected agreement against a human-human ceiling. Precision is checked separately on each cohort, to rule out a judge that favours one side. The rates are precision-adjusted. The authors also report where the instrument fails: on the three-way label, κ drops to 0.60 human-human and 0.46 judge-human, so every headline claim uses the binary label only. One disclosure the paper does not make: the judge is an Anthropic model scoring a population that includes Claude Code PRs. With two verified Claude Code pairs in the whole study, it cannot matter here.

The rates (Tables 1–2, reconciled against pdftotext)#

On each cohort's own window, 4,505 agent merges have a full 30-day window: 22.9% draw a candidate fix, 4.5% a verified one (203 PRs), 4.1% after precision adjustment. By agent (verified): Codex 5.5% (n = 2,049), Cursor 5.2% (153), Devin 3.7% (1,519), Copilot 3.5% (721), Claude Code 3.2% (63). The authors never compare the cohorts on these own-window rates. Cutting the cohorts to a shared window shrinks some agents unevenly (Cursor keeps four observable merges, Copilot has no verified fix), so the cross-cohort claim uses the shared window only:

Cohort (shared window, cut 2025-06-28)PRsCandidateVerifiedAdjusted
AI agents2,01220.9%3.68%3.34%
Humans4,06314.7%2.34%2.10%
Within-repo MH odds ratio (218 strata)1.41 [1.19–1.68]1.62 [1.10–2.39]—

The 203 verified PRs in Table 1 and the 263 verified pairs in Table 3 are different units: a merge can draw more than one fix.

Most of the excess arrives on the day of the merge. This reading comes from Fig. 3's Kaplan-Meier curves, read off the image, and the paper does not state it. The verified-fix curve for agents opens at about 2.2% on day 0, against about 0.5% for humans. At day 30 it ends near 4.5% against 2.6%. So roughly 1.7 of the final ~1.9-point gap is already there on the first day. The paper says only that "the gap is present from the first day" and that half the 30-day incidence accumulates within a week. A same-day fix PR on an agent merge, most often from the same agent, looks less like a latent defect surfacing in production and more like a unit of work the agent's workflow split across two PRs. The paper does not separate these two readings. If the second is right, some of the 1.62 measures how agents slice work, not how often agent code is wrong. (Caution on the figure: its legend reads agent n = 4,396 and human n = 5,044, which matches neither Table 1's 4,505 nor Table 2's 2,012/4,063. The caption says every merge enters with its own observation time, which accounts for part of the difference but not all of it.)

Who fixes: ownership moves to the product#

P's authorVerified pairsSelf-agentHumanOther agent
OpenAI Codex12559.2%38.4%2.4%
Devin8872.7%23.9%3.4%
GitHub Copilot4095.0%5.0%0.0%
Cursor887.5%0.0%12.5%
Claude Code20.0%50.0%50.0%
All agents26369.6%27.4%3.0%
Human110—89.1%10.9%

(Table 3, cross-checked against the PDF at ingest.) The Cursor and Claude Code rows are too small to read.

The paper frames this as code ownership with a new owner, and the ownership literature it cites is the reason to care. Bird et al. (2011) found that minor contributors predict failures better than size or churn, which is why changes are routed to owners. The paper reads its result as partly answering the worry that agent code goes unmaintained (Sawada et al. 2026 found humans author 83.2% of commits to 508 agent-generated files). It then names the gap that remains: when a human owner leaves, experienced insiders adopt the code. An agent that authors a merge and then fixes it leaves no insider behind. If the model that owns the code is deprecated, nobody on the team has accumulated experience of that code.

Three reasons to read "self-fix" narrowly.

  • "Same agent" means the same product, not the same session or the same invoking developer. The classifier asks only whether Q was opened by the agent product that opened P. In a repository that has standardised on Copilot, any fix is likely to come through Copilot, whatever the ownership relation. The paper reports no null model, such as the share of all fix PRs in the repository that come from that product. So 95% Copilot self-fixing may measure repository tool monoculture as much as ownership. The human row has the same exposure: 89.1% human-fixed is compared with nothing like the share of human fix PRs overall.
  • The pre-agent comparison mixes units. The paper calls its 89% human-fixes-human result "in line with" studies in which developers fix about half of their own defects (55.7%, 46%, about half). Those are individual-level self-fix rates. This is a class-level rate: "a human" fixed "a human's" merge. The class-level figure cannot be lower than the individual one, so the two do not test the same thing.
  • Codex is the row most exposed to misclassification. Codex commits are authored under the developer's own GitHub login, and fewer than 2% carry a marker. Its 38.4% human-fix share, the highest of any agent with more than two verified pairs, is therefore the one most likely to include the invoking developer finishing the job, or unmarked agent work counted as human. The paper's internal-validity section notes that unmarked agent PRs inflate the human-fixed share.

Merge-time signals: the pooled view and the within-repo view disagree#

RQ4 compares agent merges that later drew a fix with those that did not, on four signals. Pooled (Mann-Whitney, Cliff's δ), the groups are nearly identical. Only two differences reach significance, both small and both negative: merges that later needed a fix had less time under review (δ = −0.221) and fewer review items (δ = −0.239). On that view, lighter review goes with more fixes.

Within repositories, using conditional logistic regression that controls for files shipped, with each signal on a log₁₀ scale, the picture changes. A merge with ten times the commits carries 6.1× the odds of a verified fix (p < 0.001; Fig. 5's interval runs roughly 3× to 11×). Ten times the review items carries 2.4× (p = 0.022). Churn carries 1.3× (p < 0.001). Time under review is not significant (p = 0.18). Within a repository, then, more review activity goes with more later fixes, the opposite sign from the pooled view. The pooled "less review, more fixes" association is a between-repository effect. The authors' reading: "the PRs that needed many rounds of work before the merge are the ones most likely to need more work after it."

That undercuts the paper's own discussion heading, "code review is needed". The recommendation that verification requires reading the code rests on Rahman & Shihab's content-level model (AUC 0.671 for predicting post-merge modification), not on this paper's data. Its own data show review volume within a repo marking trouble rather than preventing it. That is consistent with review activity being a symptom of a hard change, and says nothing either way about whether review works. (See Review as the Control Point, whose review-depth edges this bears on without settling.)

Against the revert studies: opposite sign, same corpus family#

Kraishan (arXiv 2609.17598) uses AIDev PRs from the same December 2024 – July 2025 window, with a same-repo human baseline, and finds agent merges reverted less than human ones: pooled OR ≈ 0.64, and Codex at OR 0.50, the best of any vendor. This paper finds agent merges fixed more: OR 1.62. Codex has the highest verified-fix rate (5.5%) and the largest human-fix share of any agent with more than two verified pairs.

Resolution: they count different events, and each is blind to the other's. Kraishan detects reverts by commit message on the PR's most-changed file within 90 days, and states that this "misses silent rewrites". A forward fix is exactly a silent rewrite. This paper counts forward fixes that modify P's added lines on a shared, non-boilerplate file within 30 days. A plain revert would count here only if it also carried a fix-type task tag. The two results combine into one picture. Agent code is undone by revert less often than human code, and repaired by forward fix more often. The Codex pattern (few reverts, many fixes, humans doing a large share of the fixing) is the clearest case of the gap between the two channels.

Differences in population also keep the two papers from sharing a denominator. Kraishan's corpus is >100-star repositories, with 4,027 human PRs capped per repo across 810 shared repositories and 90-day windows. This one is ≥500-star, merge-only, with 5,044 human PRs from AIDev's curated table and 218 shared-repo strata on a window cut at 2025-06-28. Neither supersedes the other: both are empirical and measure different dependent variables. What this does qualify is any summary saying that controlled measurements put agent code at or below human parity on post-merge failure. That holds for reverts. It does not hold for fixes.

Scope and evidence#

  • Population: five AIDev agents, ≥500-star public repositories, 2025 merges. The authors restrict their claims to this population.
  • Effect size: a small absolute gap (about 1.3 points on the shared window) with a CI whose lower bound is 1.10.
  • Commit attribution is a lower bound: an unmarked agent commit counts as human. Codex is excluded from every commit-level figure, and commit fetches are capped at 30 per PR (under 2% of PRs reach the cap).
  • Claude Code has 130 PRs in total and only 63 with a full window (67 merged after the cutoff), and 2 verified pairs. No per-vendor conclusion about Claude Code is supported.
  • Funding and COI: Australian Composites Manufacturing CRC (ForgeX.ai grant), no declared conflict. The first author has previously published on human-in-the-loop agents at Atlassian.

Connections#

  • Agent-Vendor Heterogeneity — the opposite-sign sibling on the same corpus family (reverts below parity, fixes above it). Its open question about silent rewrites is the one this paper partly answers
  • Acceleration Whiplash — the post-deployment half of that page previously rested on two controlled sources (Google, Kraishan) putting agent code at or below parity. This is the first controlled source above parity, on forward fixes, and the gap is small (OR 1.62, lower bound 1.10), not Faros-sized
  • Efficiency Debt of AI-Generated Code — Google's ~0.9× revert ratio is on the same revert channel. This paper suggests revert ratios alone understate how often AI code needs post-merge work
  • Open Source Under Agent Contributions — the post-merge-defect half of that page's open question, measured with a verified fix label on the same kind of population, with agents 1.62× humans
  • AI as Primary Author — the authorship/responsibility gap as an ownership result: the agent fixes its own merges, so fixing does not rebuild the human insider knowledge that code ownership has depended on
  • Agentic Technical Debt — merges that took many commits are the ones that draw fixes (6.1× per tenfold commits within a repo). That is the per-PR form of an agent working without the context it needed
  • Review as the Control Point — RQ4's sign flip between the pooled and within-repo views is a warning for any review-depth-to-quality edge estimated across repositories
  • Agent-Generated Test Quality — another AIDev cut with a human baseline built from AIDev's curated human table and repository strata. It measures no test variable, so it does not answer that page's missing-baseline question, but it shows a second way to build one
  • LLM-Judge Validation — a worked example of the validation protocol: a human-human κ ceiling, cohort-split precision checks, precision-adjusted rates, and headline claims restricted to the label granularity where κ holds (binary 0.78 against three-way 0.46)
  • Verification as the New Bottleneck — the fix work lands on the agent, not the reviewer. That moves some of the post-merge verification load off humans without reducing how often it is needed
  • What the Agent-PR Oversight Numbers Can and Cannot Say — that synthesis treats survival with revert-by-message as an upper bound. This paper supplies the forward-fix side of the same bound
  • Reviewer Habituation on Agent Pull Requests — the same AIDev family's over-time approval drift, which has no outcome attached. This page's verified-fix construct is the natural outcome to join to it, to test whether late-exposure approvals are fixed forward more often than early ones

Open Questions#

  • Is agent self-fixing ownership, or repository tool monoculture? The 69.6% (and Copilot's 95%) has no null model. The test is the share of all fix PRs in the same repository and window that come from that agent product, compared against its share of fixes to its own merges. The replication package's linked pairs are enough to compute it.
  • Is the agent excess mostly same-day follow-up, rather than defects that surface later? Fig. 3 puts about 1.7 of the ~1.9-point verified gap on day 0. A split of verified fixes by time-to-fix (same day and same author vs later) would show whether the 1.62 measures latent defects or agent work-splitting.
  • Does the 1.62 survive task-type conditioning? Agents are assigned bounded tasks (the self-selection Kraishan names), and fix-typed follow-ups may be likelier on some task types. The within-repo stratification does not control task mix.

Sources#

  • Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests — Wannita Takerngsaksiri, Nhat Duong & Scott Barnett (Applied AI Initiatives, Deakin University), Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests, arXiv 2609.26847, v1 2026-09-22 (body from v2), 21pp, empirical. Cited: §3.2 (AIDev-pop cohorts), §4.1 + Tables 1–2 + Fig. 3 (candidate/verified rates, MH odds ratio, Kaplan-Meier incidence), §4.2 + Table 3 (fixer class), §4.3 + Tables 4–5 (commit author-class), §4.4 + Table 6 + Fig. 5 (merge-time signals, pooled and within-repo), §5 (ownership discussion), §6 (threats). Parse status: clean for every table cited. Tables 1, 2, 4 and 5 were reconciled cell by cell against pdftotext -layout on the local PDF at compile, and all match. Tables 3 and 6 were checked at ingest. The table-collapse flag on Table 7 is a false positive (semicolon-separated rule lists), and Table 7 is not cited. Figs. 3 and 5 were viewed. The day-0 incidence reading and Fig. 5's approximate interval bounds are read off the images, not stated in the text. Fig. 3's legend n (agent 4,396) matches no table count, a source-internal inconsistency and not a parse artefact
§ end
Cited by 12
Related articles
  • Security Debt of Agent-Generated Code

    Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Agent Review Comment Resolution

    Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex…

  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…

  • Agent-Vendor Heterogeneity

    Kraishan (Texas Tech, arXiv 2609.17598): 37,623 provenance-labelled PRs across 2,807 repos with a same-repo human basel…