資料來源#
摘要#
OpenHands 是一個開源 coding-agent 平台,也是維護並銷售此平台的公司。它最初名為 OpenDevin——Wang et al. 的論文 「OpenHands: An Open Platform for AI Software Developers as Generalist Agents」,arXiv 2407.16741,ICLR 2025——在本 wiki 的多項研究來源中都以該論文為名被引用。不同於 Claude Code 和 Codex,它的原始碼公開,因此也成為本語料庫預設的第三方 agent 框架:論文需要使用自有框架以外的 harness,讓模型執行 SWE-Bench 或 Terminal-Bench 時,通常會採用 OpenHands。
四個儲存庫組成的平台#
OpenHands 在 2026 年 7 月發布的工程文章(Coding Agents and Technical Debt),是本語料庫對 coding-agent 程式碼庫最詳盡的公開盤點,原因正是各部分分屬不同儲存庫:
| 儲存庫 | 職責 | 程式碼規模 | 12 個月內合併的 PR |
|---|---|---|---|
OpenHands/OpenHands | Agent 應用程式與伺服器 | 40.4 萬行 | 2,600 |
OpenHands/software-agent-sdk | Agent 執行環境(「Software Agent SDK」) | 33.3 萬行 | 2,036 |
OpenHands/agent-canvas | 與 agents 協作的 UI | 24.6 萬行 | 719 |
OpenHands/OpenHands-CLI | 終端機介面 | 6.7 萬行 | 324 |
截至 2026-07-08 的十二個月內,總計約 105 萬行程式碼、合併 5,679 個 PR,其中 1,778 個是錯誤修正(31%);僅 app 儲存庫就有 179 位貢獻者。Shared Harness, Differentiated Surfaces 描述的 OpenAI 架構,也採用了相同的執行環境/介面分工,只是 OpenHands 將其分別發布為四項成果。
商業定位。 OpenHands 銷售維護中的 agent,其公開主張是:應租用執行環境,而非自行分支;依序透過 prompts、MCP servers、skills/plugins 或 SDK 進行客製化。詳見 Harness Build-vs-Buy,該文進一步討論這項主張及其證據限制。
在此語料庫中的出現方式#
它大多是他人實驗所用的工具,這也提供了觀察其地位的獨立訊號:
- Single-Rollout Optimization — SAO 的 SWE-Bench Verified 結果是在 OpenHands 框架中執行(Qwen3-30B-A3B backbone、300 turns、128k context)
- Orchestration-Plan Simulation — OrchBench 用於跨框架驗證的四個實際執行框架之一,其他還有 Claude Code、SWE-mini 和 Crush
- Knowledge-Centric Self-Improvement — 報告中的 Terminal-Bench 2 比較系統(Haiku 4.5 項目表中為 13.9%)
- Agent Self-Poisoning (the CREATE-Path) — EvoMal(arXiv 2608.25776,
empirical)用於重跑技能庫自我污染攻擊的三個框架之一;未經修改,執行預設的 CodeActAgent,透過 litellm proxy 使用 DeepSeek-V4-Pro,並採用相同的預先計算檢索快取。其通用攻擊者成功率為 22.2%,低於 mini-SWE-agent 的 41.8%;但針對特定任務類別後,pytest 上升至 66.7%——與研究框架相同的上限——而反向提示則使其降至 0.0% - 在
raw/swe-pruner-pro和raw/ai-code-review-practitioner-discourse中,也被引用為基線或 agent 供應商標籤
相關連結#
- Harness Build-vs-Buy — OpenHands 自身的自建與採購主張,以及背後的十二個月 GitHub 數據;該頁說明了供應商利益關係的限制
- Shared Harness, Differentiated Surfaces — app/SDK/Canvas/CLI 的拆分,將執行環境加介面架構分別發布為不同儲存庫
- Codex — OpenHands 以其為基準比較的 OpenAI harness(合併 7,688 個 PR,約 132 萬行程式碼)
- Hermes Agent — 同一比較中的另一個開源 agent(合併 7,736 個 PR,約 175 萬行程式碼)
- Claude Code — 閉源同類產品,因這個原因未納入 OpenHands 的比較
- Orchestration-Plan Simulation — OrchBench 用來驗證其模擬器的四個實際 harness 之一
- Single-Rollout Optimization — SAO coding-agent RL 結果產生時所用的框架
- Knowledge-Centric Self-Improvement — Terminal-Bench 2 比較系統
- Agent Self-Poisoning (the CREATE-Path) — CREATE-path 攻擊可在其未經修改的預設 CodeActAgent 中重現,這是論文用來佐證自我污染源自檢索、撰寫、持久化迴圈,而非單一研究 harness 的證據之一
資料來源#
- Coding Agents and Technical Debt — Rajiv Shah (OpenHands),2026-07-28(
case-study):四個儲存庫的拆分、十二個月活動數據,以及分支/自建與採購主張 - 在以下研究中作為框架或基線被引用:Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning、OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation、Knowledge-Centric Self-Improvement、SWE-Pruner Pro: The Coder LLM Already Knows What to Prune、3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
- EVOMAL: Self-Poisoning in Self-Evolving Coding Agents — Wu、Shi、Q. Li、Zhao、X. Li、Adams、Hassan 與 Ni(Queen's University),arXiv 2608.25776,2026-08-26(
empirical):§7 與 App. C.3(Figure 5)——OpenHands 是攻擊重跑所用的兩個正式環境框架之一,並列出各任務類別的 CREATE-path 成功率與反向提示結果
Cited by 11
- Harness Build-vs-Buy×3
Every organization that decides it needs "our own coding agent" is making a make-or-buy decision,…
- Agent Self-Poisoning (the CREATE-Path)×2
Scaffolds (Figure 5, labels printed on the chart). Rerun unchanged on two production coding agents,…
- Codex×2
The one third-party accounting of Codex-the-repository in this corpus comes from a competitor:…
- ExecCritic: Learn to Test, Test to Improve×2
ExecCritic's §6 places it in a spectrum of repository-repair systems by whether validation evidence…
- Hermes Agent×2
OpenHands' twelve-month GitHub analysis (openhands coding agents technical debt, case-study,…
- Shared Harness, Differentiated Surfaces×2
Openhands — the third vendor whose four public repos make the platform split legible in line counts
- Harness Shrinkage as Models Improve
Every measurement above is taken on the system prompt. OpenHands' July 2026 GitHub analysis…
- Knowledge-Centric Self-Improvement
Highest solve rate and lowest cost in every cell — SWE-bench Pro at roughly a third of DGM's spend.…
- Entities — People, Orgs, Tools & Projects
Openhands — Open-source coding-agent platform (formerly OpenDevin, Wang et al., ICLR 2025) and the…
- Orchestration-Plan Simulation
Cross-framework validation (Figure 3) is reported as a robustness check — OrchBench correlates 0.82…
- Single-Rollout Optimization
On coding, SWE-Bench Verified (Qwen3-30B-A3B backbone, OpenHands scaffold, 300 turns, 128k…
Related articles
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Harness Build-vs-Buy
The measured price of owning a coding agent: 12 months of public GitHub activity across four harnesses (OpenHands, Code…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Skill Lift
NVIDIA SkillEvaluator's with/without-skill ablation turned into a publication gate: three pre-publication tiers (safety…
