H
Howardism
Plate IIAgent Systems中文HOWARDISM

Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves

Arena.ai's 21-combination study (7 models × Claude Code / Codex CLI / Pi, SWE-bench Lite + Terminal-Bench 2.0): harness choice moves success rate ±2–5% but multiplies cost up to 5×, geometric-mean 2.0× (Claude Code vs Pi, SWE-bench Lite); a 4-tool open-source harness (Pi) reaches the Pareto frontier on both benchmarks; and an alternative harness beats the model's own provider's harness in 9 of 12 head-to-head comparisons

Article metadata
Publication details
Published:September 25, 2026
Filed:Concept
Domain:Agent Systems
Reading:10 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves

Sources#

Summary#

Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica & Matei Zaharia, HarnessTax: How Much Does the Harness Matter for Coding Agents? (Arena.ai blog, published 2026-09-16, updated 2026-09-18, empirical). The on-page byline reads "Arena Team"; the six-author list above is recovered from the page's own BibTeX block, not stated in the visible byline. The page does not state a university affiliation for the authors — Stoica and Zaharia are widely known as UC Berkeley RISELab / Databricks-adjacent researchers, and the Acknowledgments' sponsor list (Anyscale, Broadcom, Intel, Samsung SDS…) matches the pattern of a Berkeley systems lab, but neither is a page-stated fact, so it is not asserted here as page-verified.

21 model×harness combinations (7 models: Claude Fable 5, Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna, Kimi K3 — × 3 harnesses: Claude Code, Codex CLI, and Pi, a minimal open-source harness) run on SWE-bench Lite and Terminal-Bench 2.0, with bootstrapped 95% CIs on both success rate and cost per rollout. Three findings:

  1. Harness choice barely moves success but multiplies cost.
  2. A minimal, open-source harness (Pi) is competitive — and reaches the Pareto frontier on both benchmarks.
  3. Models transfer across harnesses — a model's own provider's harness is not reliably its best pairing.

Finding 1 — cost multiplies, success barely moves#

Across all 21 SWE-bench Lite combinations (bootstrapped 95% CIs, transcribed from the article's live Plotly trace data):

ModelHarnessSuccess rate (95% CI)Cost/rollout (95% CI)Pareto?
Claude Fable 5Pi96.7% (91.1–100.0%)$0.666 (0.494–0.875)yes
Claude Fable 5Codex CLI96.7% (91.1–100.0%)$0.890 (0.697–1.127)no
Claude Fable 5Claude Code97.8% (93.3–100.0%)$1.329 (1.101–1.601)yes
Claude Opus 4.8Pi82.2% (71.1–92.2%)$0.473 (0.272–0.725)yes
Claude Opus 4.8Codex CLI88.9% (77.8–97.8%)$0.694 (0.491–0.946)no
Claude Opus 4.8Claude Code86.7% (75.6–95.6%)$0.976 (0.733–1.263)no
Claude Haiku 4.5Pi60.0% (43.3–75.6%)$0.374 (0.308–0.441)yes
GPT-5.6 SolPi74.4% (58.9–88.9%)$0.441 (0.342–0.553)yes
GPT-5.6 LunaPi53.3% (35.6–70.0%)$0.030 (0.024–0.037)yes
GPT-5.6 LunaCodex CLI55.6% (37.8–72.2%)$0.035 (0.028–0.044)yes
GPT-5.6 LunaClaude Code55.6% (38.9–72.2%)$0.152 (0.120–0.190)no

(Full 21-row table in the raw; the 10 rows omitted here are the non-Pareto Sonnet 4.6, Kimi K3, and remaining Claude-Code/Codex-CLI combinations, none on the frontier.)

Claude Fable 5 is the cleanest single instance: 97.8% (Claude Code) vs. 96.7% (Pi) — a 1.1-point success gap — at $1.329 vs. $0.666, essentially double the cost. Aggregated across shared models by geometric mean of cost ratios: Claude Code costs ~2.0× Pi and ~1.6× Codex CLI on SWE-bench Lite, and ~1.5× Pi on Terminal-Bench 2.0 (this last figure is a prose headline only — see Scope limits). Meanwhile the average harness effect on success rate stays within ±2% on SWE-bench Lite and ±5% on Terminal-Bench 2.0. Turns don't explain the gap either: Fable 5 averages 15.4 turns/attempt on Pi and 15.3 on Claude Code — same turn count, ~2× the spend.

Finding 2 — a 4-tool harness reaches the Pareto frontier#

Pi ships four tools: read, write, edit, bash. It appears on the Pareto frontier for 5 of the 7 Pareto-frontier points identified in the table above (Fable 5, Opus 4.8, Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna all reach the frontier via Pi; Claude Code and Codex CLI each contribute one frontier point). This is direct empirical corroboration, from a benchmark study external to Anthropic, of the harness-shrinkage thesis's converse: a stripped-down harness does not cost success against a 23-tool proprietary one.

The tax begins at the first model call#

Mean context in the first main model call, SWE-bench Lite, across all 7 models (Figure 3):

HarnessTools availableTool-schema size (chars)System instructions (chars)First-call context (tokens)
Pi42,873 ± 172,5471,972 ± 553
Codex CLI7.4 ± 2.518,100 ± 1,50023,500 ± 1,00011,300 ± 2,700
Claude Code2377,000 ± 7,20013,500 ± 9,40027,000 ± 6,900

By this data, Claude Code's mean first-call context (27.0k tokens) is ~13.7× Pi's (1,972 tokens) — consistent with, though sharper than, the article's rounded prose claim of "over 10×"; per-model ratios range wider still (GPT-5.6 Luna: 18.5k vs. 1,306 ≈ 14.2×; Sonnet 4.6: 32.6k vs. 2,175 ≈ 15×). Codex CLI's "7.4 ± 2.5" tool-count mean is not measurement noise — it reflects a hard split (9 tools for Claude/Kimi models and GPT-5.6 Sol, 3 for GPT-5.6 Luna).

Finding 3 — models transfer across harnesses#

Providers sometimes optimize a model for their own harness (OpenAI states GPT-5-Codex is tuned for Codex). The data does not bear this out as a general rule: across the six Anthropic and OpenAI models and both benchmarks, an alternative harness achieves the highest observed success rate in 9 of 12 comparisons. Two named instances: Sonnet 4.6 solves 68.9% of attempts in Codex CLI vs. 66.7% in Claude Code on SWE-bench Lite, at similar cost; GPT-5.6 Sol on Terminal-Bench 2.0 reaches 83.3% in Pi vs. 78.9% in Codex CLI, at about half the cost ($0.42 vs. $0.76). This is the same model-vs-scaffold decoupling Measuring Beyond Accuracy Saturation documents on CORE-Bench (scaffolds disagreeing on 31% of tied-accuracy capsules) — here on the cross-vendor axis rather than the open-vs-proprietary one.

Scope limits and evidence caveats#

  • Both benchmarks are open-source and long-published (SWE-bench Lite, Terminal-Bench 2.0) — the authors flag possible training-data contamination themselves; results may not generalize to unseen or private benchmarks.
  • Terminal-Bench 2.0 per-combination numbers were not captured for this raw. All four data figures are Plotly charts mounted client-side inside sandboxed, cross-origin srcdoc iframes; a Playwright pass read the live .data trace off each chart via per-frame evaluate(), but clicking each figure's "Terminal-Bench 2" tab visibly toggled the UI while the underlying trace ids stayed prefixed swe:... — a site-side toggle bug, not a fetch failure. Only the article's prose Terminal-Bench 2.0 headlines (the ~1.5× Pi cost ratio, the ±5% success band, the 83.3%-vs-78.9% GPT-5.6 Sol comparison) are carried into this page; no Terminal-Bench 2.0 table exists in the raw or here. Treat any Terminal-Bench 2.0 number above as prose-sourced, not table-reconciled.
  • Vendor-adjacent voice. Arena.ai is a commercial LLM-evaluation company (the LMArena lineage); the piece doubles as a showcase for Arena's own profiling/trace-capture methodology and closes with a promise to "publicly release profiling traces." Rated empirical here on the strength of the bootstrapped CIs and the real, benchmark-measured cost/success grid — this is genuine controlled measurement, not an announcement — but it is not third-party-audited or independently replicated, and Arena has no product stake in Claude Code, Codex CLI, or Pi specifically, which limits (but does not eliminate) the COI concern relative to a harness vendor grading itself.

Connections#

  • Cost-per-Task Over Cost-per-Token — the cost-axis instance this page supplies: a fourth data point (after Cursor's model-mix, Writer's harness-outweighs-model-menu, and Databricks' open-weight bench) showing the same model's cost varying up to 2× by harness alone, at a success delta within noise
  • Harness Shrinkage as Models Improve (hub) — external corroboration, from a benchmark study outside Anthropic, that a stripped 4-tool harness does not cost success against a 23-tool proprietary one; the ~13.7× first-call-context gap is a second-vendor measurement of the same context-bloat problem Cat Wu's pruning discipline targets
  • Compute-Controlled Benchmarking — harness choice is an undisclosed budget surface: two agents scoring within a few points of each other on SWE-bench Lite can differ 2–5× in dollar cost purely by harness, the same omitted-variable problem this page's grid names for the compute axis
  • Measuring Beyond Accuracy Saturation — corroborates "Decoupling model and scaffold" on a new axis: 9 of 12 cross-vendor comparisons favor a model's non-native harness, extending CORE-Bench's within-vendor scaffold-disagreement finding to the cross-provider case
  • Agent Harness Engineering — a controlled, cross-model complement to Cline's single-vendor production anecdote ("a 2024-shaped harness taxes a 2026 model"): here the tax is measured as a cost multiplier, across 7 models and 3 harnesses, holding the benchmark fixed
  • Claude Code — the highest-cost harness in every shared-model comparison in this study; the entity page does not currently carry this cost comparison

Open Questions#

  • Does the harness-cost gap (2–5×, ±2–5% success) hold on benchmarks the tested models have not seen in training, or is it an artifact of SWE-bench Lite / Terminal-Bench 2.0 being long-public?
  • The paper proposes a harness that "adapts as tasks unfold while remaining general," replacing the user's harness-selection decision — does any deployed system attempt per-task harness routing, and does it recover most of the tax or just relocate the selection problem?

Sources#

  • HarnessTax: How Much Does the Harness Matter for Coding Agents? — Pan, Yang, Arabzadeh, Chiang, Stoica & Zaharia, Arena.ai blog, published 2026-09-16 / updated 2026-09-18, empirical. Fetched via r.jina.ai extraction (plain WebFetch returns a truncated Next.js shell); on-page metadata and BibTeX recovered via direct curl. Capture warning: all four data figures are live-mounted Plotly charts read via Playwright's per-iframe frame.evaluate(), not screenshot transcription — reliable for SWE-bench Lite (all 21 combinations, confirmed by two independently rendered figures agreeing), but Terminal-Bench 2.0's per-combination data was not recoverable (site-side tab-toggle bug); only prose headlines for Terminal-Bench 2.0 are used anywhere in this page.
§ end
Cited by 7
Related articles