Sources#
Summary#
Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica & Matei Zaharia, HarnessTax: How Much Does the Harness Matter for Coding Agents? (Arena.ai blog, published 2026-09-16, updated 2026-09-18, empirical). The on-page byline reads "Arena Team"; the six-author list above is recovered from the page's own BibTeX block, not stated in the visible byline. The page does not state a university affiliation for the authors — Stoica and Zaharia are widely known as UC Berkeley RISELab / Databricks-adjacent researchers, and the Acknowledgments' sponsor list (Anyscale, Broadcom, Intel, Samsung SDS…) matches the pattern of a Berkeley systems lab, but neither is a page-stated fact, so it is not asserted here as page-verified.
21 model×harness combinations (7 models: Claude Fable 5, Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna, Kimi K3 — × 3 harnesses: Claude Code, Codex CLI, and Pi, a minimal open-source harness) run on SWE-bench Lite and Terminal-Bench 2.0, with bootstrapped 95% CIs on both success rate and cost per rollout. Three findings:
- Harness choice barely moves success but multiplies cost.
- A minimal, open-source harness (Pi) is competitive — and reaches the Pareto frontier on both benchmarks.
- Models transfer across harnesses — a model's own provider's harness is not reliably its best pairing.
Finding 1 — cost multiplies, success barely moves#
Across all 21 SWE-bench Lite combinations (bootstrapped 95% CIs, transcribed from the article's live Plotly trace data):
| Model | Harness | Success rate (95% CI) | Cost/rollout (95% CI) | Pareto? |
|---|---|---|---|---|
| Claude Fable 5 | Pi | 96.7% (91.1–100.0%) | $0.666 (0.494–0.875) | yes |
| Claude Fable 5 | Codex CLI | 96.7% (91.1–100.0%) | $0.890 (0.697–1.127) | no |
| Claude Fable 5 | Claude Code | 97.8% (93.3–100.0%) | $1.329 (1.101–1.601) | yes |
| Claude Opus 4.8 | Pi | 82.2% (71.1–92.2%) | $0.473 (0.272–0.725) | yes |
| Claude Opus 4.8 | Codex CLI | 88.9% (77.8–97.8%) | $0.694 (0.491–0.946) | no |
| Claude Opus 4.8 | Claude Code | 86.7% (75.6–95.6%) | $0.976 (0.733–1.263) | no |
| Claude Haiku 4.5 | Pi | 60.0% (43.3–75.6%) | $0.374 (0.308–0.441) | yes |
| GPT-5.6 Sol | Pi | 74.4% (58.9–88.9%) | $0.441 (0.342–0.553) | yes |
| GPT-5.6 Luna | Pi | 53.3% (35.6–70.0%) | $0.030 (0.024–0.037) | yes |
| GPT-5.6 Luna | Codex CLI | 55.6% (37.8–72.2%) | $0.035 (0.028–0.044) | yes |
| GPT-5.6 Luna | Claude Code | 55.6% (38.9–72.2%) | $0.152 (0.120–0.190) | no |
(Full 21-row table in the raw; the 10 rows omitted here are the non-Pareto Sonnet 4.6, Kimi K3, and remaining Claude-Code/Codex-CLI combinations, none on the frontier.)
Claude Fable 5 is the cleanest single instance: 97.8% (Claude Code) vs. 96.7% (Pi) — a 1.1-point success gap — at $1.329 vs. $0.666, essentially double the cost. Aggregated across shared models by geometric mean of cost ratios: Claude Code costs ~2.0× Pi and ~1.6× Codex CLI on SWE-bench Lite, and ~1.5× Pi on Terminal-Bench 2.0 (this last figure is a prose headline only — see Scope limits). Meanwhile the average harness effect on success rate stays within ±2% on SWE-bench Lite and ±5% on Terminal-Bench 2.0. Turns don't explain the gap either: Fable 5 averages 15.4 turns/attempt on Pi and 15.3 on Claude Code — same turn count, ~2× the spend.
Finding 2 — a 4-tool harness reaches the Pareto frontier#
Pi ships four tools: read, write, edit, bash. It appears on the Pareto frontier for 5 of the 7 Pareto-frontier points identified in the table above (Fable 5, Opus 4.8, Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna all reach the frontier via Pi; Claude Code and Codex CLI each contribute one frontier point). This is direct empirical corroboration, from a benchmark study external to Anthropic, of the harness-shrinkage thesis's converse: a stripped-down harness does not cost success against a 23-tool proprietary one.
The tax begins at the first model call#
Mean context in the first main model call, SWE-bench Lite, across all 7 models (Figure 3):
| Harness | Tools available | Tool-schema size (chars) | System instructions (chars) | First-call context (tokens) |
|---|---|---|---|---|
| Pi | 4 | 2,873 ± 17 | 2,547 | 1,972 ± 553 |
| Codex CLI | 7.4 ± 2.5 | 18,100 ± 1,500 | 23,500 ± 1,000 | 11,300 ± 2,700 |
| Claude Code | 23 | 77,000 ± 7,200 | 13,500 ± 9,400 | 27,000 ± 6,900 |
By this data, Claude Code's mean first-call context (27.0k tokens) is ~13.7× Pi's (1,972 tokens) — consistent with, though sharper than, the article's rounded prose claim of "over 10×"; per-model ratios range wider still (GPT-5.6 Luna: 18.5k vs. 1,306 ≈ 14.2×; Sonnet 4.6: 32.6k vs. 2,175 ≈ 15×). Codex CLI's "7.4 ± 2.5" tool-count mean is not measurement noise — it reflects a hard split (9 tools for Claude/Kimi models and GPT-5.6 Sol, 3 for GPT-5.6 Luna).
Finding 3 — models transfer across harnesses#
Providers sometimes optimize a model for their own harness (OpenAI states GPT-5-Codex is tuned for Codex). The data does not bear this out as a general rule: across the six Anthropic and OpenAI models and both benchmarks, an alternative harness achieves the highest observed success rate in 9 of 12 comparisons. Two named instances: Sonnet 4.6 solves 68.9% of attempts in Codex CLI vs. 66.7% in Claude Code on SWE-bench Lite, at similar cost; GPT-5.6 Sol on Terminal-Bench 2.0 reaches 83.3% in Pi vs. 78.9% in Codex CLI, at about half the cost ($0.42 vs. $0.76). This is the same model-vs-scaffold decoupling Measuring Beyond Accuracy Saturation documents on CORE-Bench (scaffolds disagreeing on 31% of tied-accuracy capsules) — here on the cross-vendor axis rather than the open-vs-proprietary one.
Scope limits and evidence caveats#
- Both benchmarks are open-source and long-published (SWE-bench Lite, Terminal-Bench 2.0) — the authors flag possible training-data contamination themselves; results may not generalize to unseen or private benchmarks.
- Terminal-Bench 2.0 per-combination numbers were not captured for this raw. All four data figures are Plotly charts mounted client-side inside sandboxed, cross-origin
srcdociframes; a Playwright pass read the live.datatrace off each chart via per-frameevaluate(), but clicking each figure's "Terminal-Bench 2" tab visibly toggled the UI while the underlying traceidsstayed prefixedswe:...— a site-side toggle bug, not a fetch failure. Only the article's prose Terminal-Bench 2.0 headlines (the ~1.5× Pi cost ratio, the ±5% success band, the 83.3%-vs-78.9% GPT-5.6 Sol comparison) are carried into this page; no Terminal-Bench 2.0 table exists in the raw or here. Treat any Terminal-Bench 2.0 number above as prose-sourced, not table-reconciled. - Vendor-adjacent voice. Arena.ai is a commercial LLM-evaluation company (the LMArena lineage); the piece doubles as a showcase for Arena's own profiling/trace-capture methodology and closes with a promise to "publicly release profiling traces." Rated
empiricalhere on the strength of the bootstrapped CIs and the real, benchmark-measured cost/success grid — this is genuine controlled measurement, not an announcement — but it is not third-party-audited or independently replicated, and Arena has no product stake in Claude Code, Codex CLI, or Pi specifically, which limits (but does not eliminate) the COI concern relative to a harness vendor grading itself.
Connections#
- Cost-per-Task Over Cost-per-Token — the cost-axis instance this page supplies: a fourth data point (after Cursor's model-mix, Writer's harness-outweighs-model-menu, and Databricks' open-weight bench) showing the same model's cost varying up to 2× by harness alone, at a success delta within noise
- Harness Shrinkage as Models Improve (hub) — external corroboration, from a benchmark study outside Anthropic, that a stripped 4-tool harness does not cost success against a 23-tool proprietary one; the ~13.7× first-call-context gap is a second-vendor measurement of the same context-bloat problem Cat Wu's pruning discipline targets
- Compute-Controlled Benchmarking — harness choice is an undisclosed budget surface: two agents scoring within a few points of each other on SWE-bench Lite can differ 2–5× in dollar cost purely by harness, the same omitted-variable problem this page's grid names for the compute axis
- Measuring Beyond Accuracy Saturation — corroborates "Decoupling model and scaffold" on a new axis: 9 of 12 cross-vendor comparisons favor a model's non-native harness, extending CORE-Bench's within-vendor scaffold-disagreement finding to the cross-provider case
- Agent Harness Engineering — a controlled, cross-model complement to Cline's single-vendor production anecdote ("a 2024-shaped harness taxes a 2026 model"): here the tax is measured as a cost multiplier, across 7 models and 3 harnesses, holding the benchmark fixed
- Claude Code — the highest-cost harness in every shared-model comparison in this study; the entity page does not currently carry this cost comparison
Open Questions#
- Does the harness-cost gap (2–5×, ±2–5% success) hold on benchmarks the tested models have not seen in training, or is it an artifact of SWE-bench Lite / Terminal-Bench 2.0 being long-public?
- The paper proposes a harness that "adapts as tasks unfold while remaining general," replacing the user's harness-selection decision — does any deployed system attempt per-task harness routing, and does it recover most of the tax or just relocate the selection problem?
Sources#
- HarnessTax: How Much Does the Harness Matter for Coding Agents? — Pan, Yang, Arabzadeh, Chiang, Stoica & Zaharia, Arena.ai blog, published 2026-09-16 / updated 2026-09-18,
empirical. Fetched viar.jina.aiextraction (plain WebFetch returns a truncated Next.js shell); on-page metadata and BibTeX recovered via directcurl. Capture warning: all four data figures are live-mounted Plotly charts read via Playwright's per-iframeframe.evaluate(), not screenshot transcription — reliable for SWE-bench Lite (all 21 combinations, confirmed by two independently rendered figures agreeing), but Terminal-Bench 2.0's per-combination data was not recoverable (site-side tab-toggle bug); only prose headlines for Terminal-Bench 2.0 are used anywhere in this page.
Cited by 7
- Agent Harness Engineering×3
Harness Tax — Cline's single-vendor tax measured as a cross-vendor grid: cost up to 5× apart for…
- Compute-Controlled Benchmarking×3
Kimi K3's card names the harness per score but publishes no budget for it. HarnessTax (harness tax…
- Cost-per-Task Over Cost-per-Token×3
HarnessTax runs the harness-as-cost-lever question as a controlled grid rather than a single swap:…
- Harness Shrinkage as Models Improve×3
HarnessTax (harness tax coding agent benchmarks, empirical) is a third-party benchmark rather than…
- Measuring Beyond Accuracy Saturation×3
harness tax coding agent benchmarks — Pan, Yang, Arabzadeh, Chiang, Stoica & Zaharia, Arena.ai…
- Agent Systems & Harness Engineering
Harness Tax — Arena.ai's 21-combination study (7 models × Claude Code / Codex CLI / Pi, SWE-bench…
- Open Questions Backlog
Harness Tax ×2 (oldest 4d) — Does the harness-cost gap (2–5×, ±2–5% success) hold on benchmarks the…
Related articles
- Agent-Authored Harness Optimization
An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Seven instances dis…
- Orchestration Sets Token Economics
Writer's controlled harness swap — same 22 tasks, same six models, same judges and price table, only the orchestration…
- Shared Harness, Differentiated Surfaces
OpenAI merged Codex and ChatGPT Work onto one agent harness and differentiated only the UX layer — git-state visibility…
- Cost-per-Task Over Cost-per-Token
Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…
- Claude Code Best Practices
Anthropic's guide to effective Claude Code usage: context management, verification-driven development, explore→plan→cod…
