H
Howardism
Plate IIModel Capability & Training中文HOWARDISM

ExecCritic: Learn to Test, Test to Improve

Tao, Peng, Wang et al. (UW-Madison / Microsoft Research / Georgia Tech): a test-verify-revise scaffold that separates test construction from source-code repair into two independently-trained Qwen-3.5-35B-A3B roles — a Test agent whose bundle a fail-closed harness qualifies and freezes, and a Repair agent that revises only the source patch from that fixed test's feedback. On SWE-bench Verified, holding the Repair agent fixed, tests from the untrained Test agent *cut* resolved rate from 61.2% to 57.3% while GPT-5.6-generated tests raise it to 65.3% — feedback helps or hurts depending on test quality, not on whether feedback exists. Role-specific RL raises Test-agent Base-to-Gold success 22.2%→62.2% and Repair-agent Round-0 61.2%→68.3%; composing the two trained agents reaches 72.6%, +11.4pp over the original no-test baseline, without a stronger model or Oracle feedback at evaluation time

Article metadata
Publication details
Published:September 24, 2026
Filed:Concept
Domain:Model Capability & Training
Reading:16 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for ExecCritic: Learn to Test, Test to Improve

Sources#

Summary#

ExecCritic (Tao, Peng, Wang, Wang, Cheng, Yao, Wu, Ge, Li & Gao — University of Wisconsin-Madison / Microsoft Research / Georgia Tech, arXiv 2609.09133, 2026-09-08, empirical) trains repository-repair agents to distrust their own validation evidence. The premise: when the same trajectory writes both a patch and the test that checks it, a blind spot in one becomes a blind spot in both — the patch and the test can "agree on the same incomplete interpretation of the issue" and a wrong fix passes a wrong check. The fix is architectural before it is a training recipe: split test-writing and repair into two agents that cannot see or influence each other's artifact, then train each role separately against the outcome it actually controls. On the same execution-feedback-RL family RLEF belongs to, this is the first result in the corpus to show that the source of the feedback, not its presence, is what determines whether the loop helps.

The coupling argument#

Formally (§2): an instance is x = (d, R_B), an issue description and a buggy checkout. In a conventional agent-controlled loop, (p, E) ~ π_ω(·|x) — the Repair patch p and its validation evidence E are sampled from the same policy and trajectory. When E = ∅ the agent submits with no local check; when E ≠ ∅, a shared misinterpretation can shape the patch, the evidence that appears to validate it, and the decision to stop. The worked example: a bug triggers only on an empty list; the agent fixes the common case, writes a test that only covers the common case, the test passes, and the agent concludes the issue is resolved with the original bug intact. The official evaluator V*_x(p) ∈ {0,1} — hidden fail-to-pass and pass-to-pass tests — remains the ground truth throughout; nothing in the scaffold touches it.

ExecCritic's fix generates the test without access to the candidate Repair trajectory and holds it fixed during revision. The paper is explicit that this is independence of generation context and write permissions, not statistical independence of errors: "the two agents can still share a mistaken interpretation of the issue" (see Open Questions).

The scaffold#

Three components, all present in every run:

  1. Test agent (π^T_φ) explores the repository from (d, R_B) alone and submits a Test bundle b = (Δ_b, c_b, κ_b) — a repository-native test patch, its exact execution command, and a JSON behavior contract. Up to five submission attempts per episode.
  2. Fail-closed harness qualifies b: it must execute validly and produce a clean failure on Base (B_x(b) = 1). The first bundle to clear this bar is frozen; qualification does not require passing on Gold, so the harness is checking "does this test currently fail," not "is this test correct." If no attempt clears the Base gate, the instance is marked a Base-gate failure, feedback-guided repair never launches, and the Round-0 patch alone goes to the official evaluator — these instances stay in the resolved-rate denominator rather than being dropped.
  3. Repair agent (π^R_θ) revises source only, never the test. Each round resets an isolated workspace to R_B, applies the current patch and the fixed bundle, and returns z_t ∈ {PASS, FAIL} plus bounded execution output o_t. A local PASS ends the episode immediately (a stopping condition, not a claim of correctness); FAIL lets the agent continue revising or submit anyway. T_max = 5 revision rounds of 40 turns each; if nothing passes, the harness force-submits the last candidate.

Base-to-Gold success, the Test agent's core training and evaluation metric, is Q_x(b) = B_x(b)·G_x(b) — a clean Base failure and a pass after applying the reference (Gold) patch. Gold outcomes and labeled candidate patches feed training rewards and offline audits only; they never filter which bundles reach the Repair stage or enter the Test agent's own trajectory.

Role-specific RL#

Both roles start from Qwen-3.5-35B-A3B, trained separately, on SWE-ReBench (Badertdinov et al., 2025), evaluated on SWE-bench Verified, inside Orchard (Peng et al., 2026), an open-source agentic-RL/sandbox framework. GPT-5.6-sol (medium reasoning effort) is the stronger-model reference throughout.

Learn to Test. SFT on 5,000 chain-of-thought trajectories from a DeepSeek-V4-Flash-0731 teacher, then 200 steps of on-policy RL with DAPO-style dynamic sampling (16 nonzero-variance issue groups × 8 rollouts = 128 rollouts/update). The reward, over K sampled bundles per task, uses balanced accuracy BA_x(b) = ½(TPR + TNR) on up to eight deduplicated labeled candidate Repair patches:

  • −0.2 — no valid submission
  • 0 — valid submission, but B_x(b) = 0 or G_x(b) = 0
  • 0.2 — Q_x(b) = 1, BA < 0.8
  • 0.5 — Q_x(b) = 1, 0.8 ≤ BA < 1
  • 1.0 — Q_x(b) = 1, BA = 1

Any candidate-execution operational error zero-advantages and masks the whole trajectory rather than scoring it as a behavioral failure. A trajectory that runs more than eight turns past the shortest one in its reward bracket gets half credit.

Test to Improve. Repair training is restricted to SWE-ReBench issues where the Base model's empirical Round-0 accuracy is below 0.4 (frequently-solved issues carry no revision signal), 100 GRPO steps, same dynamic-sampling shape. One Oracle F2P test is selected and held fixed per task; the Repair agent sees only its execution feedback, not the test source. Reward for the terminal patch p_t (z^train_t the selected-test outcome, u_t the official outcome):

  • 0 — invalid terminal patch (including a forbidden test edit)
  • 0.1 — selected test FAILs
  • 0.2 — selected test PASSes, official FAILs
  • 1.0 — both PASS, t > 0 (after revision)
  • 1.5 — both PASS, t = 0 (Round-0, direct solve)

The gap between 0.2 and 1.0 is the whole point: satisfying the training test is worth an order of magnitude less than satisfying it and the full hidden suite. The larger Round-0 bonus (1.5 vs 1.0) is a deliberate counterweight against the policy learning to defer an incomplete Round-0 solution in the hope that revision will fix it later.

Results#

Headline. Composing the RL-trained Test agent (RL-35B) with the trained Repair agent reaches 72.6% on SWE-bench Verified — +4.3pp over the trained Repair agent's own no-test Round-0 (68.3%) and +11.4pp over the original untrained no-test baseline (61.2%) — without GPT-5.6 or Oracle feedback at evaluation.

Test quality determines whether feedback helps at all (Table 2, holding the Repair agent fixed). Ingest note: this is the paper's sharpest single result, and it isolates the Test source as the only variable:

Test source (feeding the Base Repair agent)Resolved (%)Δ vs no-test
None (Round 0)61.2—
Base-35B (untrained Qwen Test agent)57.3−3.9
RL-35B (trained Qwen Test agent)64.1+2.9
GPT-5.6-sol65.3+4.1
Oracle F2P69.4+8.2

An untrained Test agent's tests make the same Repair agent worse than having no test at all — a bad check can steer revision toward an incomplete or incorrect requirement rather than the real fault. The ordering is consistent with the independent Base-to-Gold measurements (Table 1: Base Qwen 22.2%, GPT-5.6-sol 87.8%) — the source that produces more Base-to-Gold-reliable tests is the source whose feedback helps repair.

Test-agent training (Table 1, Base-to-Gold success on SWE-bench Verified). SFT alone: 22.2% → 39.6% (+17.4pp). RL on top: → 62.2% (+22.6pp further), landing near Codex-5.3 (61.0%) and still 11.2pp below the SFT teacher (DeepSeek-V4-Flash-0731, 73.4%) and 25.6pp below GPT-5.6-sol (87.8%). SFT-dataset scaling (Fig. 6a) shows diminishing returns by 5K trajectories (10K adds nothing on this metric) while the teacher reference stays near 73% — RL is what closes most of the remaining gap, and only from an SFT-initialized policy: RL from the raw Base checkpoint stays low and oscillatory (no differentiated within-group reward to train on), while RL from the SFT checkpoint rises steadily.

Repair-agent training (Table 2, full grid). No-test Round-0 rises 61.2% → 68.3% (+7.1pp, direct repair capability). Test-to-Improve then adds +2.9pp to the untrained Repair agent under RL-35B feedback and +4.3pp to the trained one — different starting points, so the two gains are not a matched estimate of "feedback-use capability" in isolation, but both trained components help the system regardless of which version of the other component it is paired with. With the trained pair, Oracle F2P feedback reaches 77.6%, a 5.0pp residual gap above the generated-test composition the authors attribute to remaining headroom in test quality and coverage.

Off-the-shelf frontier agents (Fig. 5a) show the same feedback-depends-on-coverage pattern at a different scale. GPT-5.6 on SWE-bench Verified: 81.5% → 85.7% with generated tests (+4.2), → 89.7% with Oracle F2P (+8.2). On SWE-bench Pro: 61.6% → 62.3% with generated tests (+0.7 — almost nothing), → 73.4% with Oracle (+11.8). Pro's median/mean F2P test count per instance is 3/14.43 against Verified's 1/3.03 — one generated bundle targeting one issue behavior plausibly leaves the rest of a broader spec uncovered, though the paper is careful that this is a correlational read, not an isolated cause.

Ablations. The direct-solve bonus matters for the Round-0/final balance (Table 3): full T2I reaches 68.3%/72.6%, dropping the bonus trades a lower Round-0 (66.4%) for a smaller final gain (71.4%) — consistent with the bonus discouraging the policy from deferring to revision. SFT-teacher reasoning visibility matters far more after RL than after SFT alone (Table 4): a non-CoT GPT-5.6-sol teacher and the CoT DeepSeek-V4-Flash-0731 teacher differ by only 4.2pp post-SFT (35.4 vs 39.6) but by 26.0pp post-RL (36.2 vs 62.2) — though teacher identity and CoT visibility are confounded, so this isn't a clean causal isolation of chain-of-thought. Oracle-test training beats generated-test training at Repair time by a modest 0.7–1.7pp (Table 6) — cleaner training feedback helps, but not by much, since the terminal reward already trains the whole trajectory including the Round-0 patch.

Cross-language generalization is real but uneven (Table 5). Post-trained only on Python, the Test agent's Base-to-Gold success on SWE-bench Multilingual: Rust 66.7%, C++ 58.3% (both above the 62.2% Python in-language reference), C 26.1%, Java 2.4%. The method transfers the behavioral-contract-inference skill without language-specific post-training, but reliability collapses on some toolchains — these are Test-generation rates, not downstream Repair-resolution rates, so the practical repair impact of a near-2.4%-reliable Test agent on Java is not separately measured.

Repair overhead stays small. Among the 120 trajectories that entered repair with the trained pair, revision used an average of only 13 additional agent turns to go from 68.3% to 72.6% — most of the gain does not consume the full 5-round budget. Test generation is a one-time cost per issue/source, reused across the three Repair runs each resolved-rate figure averages over.

ExecCritic's §6 places it in a spectrum of repository-repair systems by whether validation evidence stays fixed during revision. Repository-level agents that interleave exploration, editing, execution and revision — SWE-agent, OpenHands, Mini-SWE-Agent — are the generic harness family this scaffold builds on top of, generating their own checks with no separation of roles. TDFlow freezes supplied or generated tests the same way ExecCritic does; SpecRover, Dynamic Cogeneration, CodeMonkeys, InfCode, Agent-CoEvo, CoHarden and TDD-Agent instead update, select, or harden tests as repair proceeds — an evolving test can correct weak evidence but also changes the acceptance criterion mid-episode, which is exactly the coupling risk ExecCritic is built to avoid. The classical automated-program-repair line (DiffTGen, Opad, Patch-Sim, RGT, Poracle) attacks the same oracle problem — a patch can be plausible against an available suite yet semantically wrong — with post-hoc filters rather than a training-time role split.

Limitations#

Test and Repair are trained as separate policies, composed only after training — different action spaces, trajectory structures and reward shapes made a shared-weight objective out of scope here. The authors name joint training, with role conditioning determining test-vs-repair behavior, as future work, alongside multi-check harnesses and stronger Test-agent weights.

Connections#

  • RL from Execution Feedback (RLEF) — the same execution-feedback-RL lineage as RLEF, one level more structured: RLEF puts test execution in the reward loop and studies feedback richness and credit granularity on CodeContests, where public/private tests come from one curated benchmark split; ExecCritic asks the question RLEF's design never separates out — whether the test itself is a valid oracle — and answers it at repository scale, where nobody hands the agent a matched public/private pair. The two are not competing answers to the same open question: RLEF's is about how much information a fixed-format check returns, ExecCritic's is about whether the check is correct at all — read the latter as a precondition that binds before the former's richness question is even reachable
  • Agent-Generated Test Quality — the observational sibling this supplies a causal, training-time complement to. Jhanglani et al. and Dipongkor et al. measure whether agent-written tests are broad and whether they cover the diff; ExecCritic runs the controlled version of the same question — hold the Repair agent fixed, swap only the Test source — and gets a signed answer: an untrained Test agent's checks make repair worse than no test (57.3% vs 61.2%), a reliable one makes it meaningfully better (65.3%, 69.4% Oracle). "Breadth is intrinsic, targeting is relational" gets a third axis here: reliability, measured as Base-to-Gold success, and it is reliability, not breadth, that determines whether a downstream repair loop benefits
  • Reward Hacking — the shared-trajectory coupling argument in §2 is a structural reward-hacking guard in the same family as RLEF's public/private test split: withhold write access to the validation artifact from the policy being evaluated, rather than trying to shape a reward the policy can read. The harness's fail-closed Base-gate is the enforcement mechanism — an unqualified test cannot enter the reward at all rather than entering it as a noisy zero
  • Verification as the New Bottleneck (hub) — a training-time instance of the same claim: once the harness makes validation independent and load-bearing, the ceiling on repair performance is bounded above by test-writing reliability (Table 1's 62.2%, well below the 87.8% GPT-5.6-sol reference), not by repair capability alone
  • OpenHands — cited as one of the repository-level agent scaffolds (with SWE-agent, Mini-SWE-Agent) whose exploration-editing-execution-revision loop ExecCritic's Repair agent specializes and constrains, rather than one it is evaluated inside

Open Questions#

  • The paper concedes independence is about generation context and write permissions, not statistical independence of errors — "the two agents can still share a mistaken interpretation of the issue." How much of the residual gap to Oracle feedback (72.6% generated vs 77.6% Oracle, both on the trained Repair agent) is correlated misinterpretation the Test and Repair agents share from the same ambiguous issue text, versus independent Test-quality or Repair-capability shortfalls? Falsifiable: compare the composition's error set against Oracle-fed Repair's error set on the same instances and classify overlap by whether the Test bundle target and the Repair agent's misreading match.
  • Cross-language Base-to-Gold success drops from 66.7% (Rust) to 2.4% (Java) with no target-language post-training. Does the gap trace to test-infrastructure discovery (finding the right build/test runner and invocation command) or to behavioral-contract inference itself (understanding what the issue requires)? The two failure modes call for different fixes — tooling scaffolding versus more post-training data — and the paper's single-generation-per-instance protocol doesn't distinguish them.

Sources#

  • ExecCritic: Learn to Test, Test to Improve for Coding Agents — Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li & Jianfeng Gao (University of Wisconsin-Madison / Microsoft Research / Georgia Tech), ExecCritic: Learn to Test, Test to Improve for Coding Agents, arXiv 2609.09133, 2026-09-08, empirical. Code at github.com/MSR-Orchard/execcritic. §2 the coupling formalism; §3 the scaffold and Learn-to-Test/Test-to-Improve mechanics (Figs. 2–4); §4 the reward definitions; §5 experiments (Tables 1–6, Fig. 5a–b, Fig. 6); §6 related work; §7 limitations. Parse notes. PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 35pp, rapidocr, confidence excellent). Results Tables 1–6 parsed clean and are quoted directly; only the appendix hyperparameter tables carry the collapse/weld warnings this vault's docling audits generally flag, and this page cites none of them. Evidence. empirical holds on a full read: controlled ablations (Tables 2, 3, 6) that vary one component while holding the other fixed, resolved rates averaged over three Repair runs per condition, a fail-closed harness whose qualification criterion is checkable independent of the training signal, and cross-benchmark generalization tests (SWE-bench Pro, SWE-bench Multilingual) run without additional post-training. One authorship note: the lead author's internship affiliation is Microsoft Research, which also employs six of the ten co-authors — an internal evaluation, not an independent audit, though the benchmark (SWE-bench Verified/Pro/Multilingual) and the reference models compared against (Codex-5.3, DeepSeek-V4-Flash, GPT-5.6-sol) are all third-party.
§ end
Cited by 5
  • Agent-Generated Test Quality×2

    Execcritic — the causal, training-time complement above: holding a Repair agent fixed and swapping…

  • RL from Execution Feedback (RLEF)

    Execcritic — the repository-scale descendant that answers the question RLEF's own binary-reward…

  • Model Capability & Training

    Execcritic — Tao, Peng, Wang et al. (UW-Madison / Microsoft Research / Georgia Tech): a…

  • Open Questions Backlog

    Execcritic ×2 (oldest 5d) — The paper concedes independence is about generation context and write…

  • Reward Hacking

    Execcritic — a structural guard against a specific hacking channel, stated as a formal argument…

Related articles