資料來源#
- Agent swarms and the new model economics
- AI Engineering Report 2026: The Acceleration Whiplash
- GhostJacking Attacks: Half of the Fortune 500 Run These Tools. Getting Blocked by the Firewall Was the Way to Take Over Their AI Agents
- Introducing Projects
- One Fake Bug Report Hijacked a $250 Billion Company's AI Agent – Then 100+ More
- The Week of Sandbox Escapes
- Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents
摘要#
這家公司打造了 Cursor IDE——一款代理式程式碼編輯器——以及自家的 Composer 模型系列,並經營公開的研究/工程部落格。在這份語料中,它是僅有的兩家(另一家是 Anthropic)曾詳細說明如何在實際工程任務中運行極大規模代理程式分流的供應商之一,也是唯一一家在說明中以相同模型與相同時間預算,刻意比較舊版與新版 harness 的供應商。
Cursor 透過三扇幾乎毫無關聯的門進入這份知識庫;評估任何來自 Cursor 的主張時,最好將這三者分開看待。
1. 蜂群工程的發布者#
Agent swarms and the new model economics(Wilson Lin,2026-07-20,case-study)是繼 Bun Zig→Rust 移植之後,語料中的第二個大型蜂群案例研究,也是兩者中工程設計更周密的一個。其組成如下:
- 兩種角色,一棵遞迴樹。 規劃者(最聰明的模型)負責拆解任務並委派工作;工作者(快速、低成本的模型)負責執行。「規劃者從不實作……工作者從不規劃。」Cursor 表示,這種設計能擴展的原因在於脈絡效率,而非平行處理——詳見 Multi-Agent Collective Intelligence。
- 從零打造的版本控制系統。 早期的瀏覽器蜂群在 Git 上每小時最多提交約 1,000 次;新系統的峰值則是每秒約 1,000 次提交。提升吞吐量並非唯一動機——每項變更都會經過 VCS,因此碰撞會最先在此顯現,多種協調機制也設於其中。
- 具名的失敗模式與具名的修正方式——腦裂設計、規劃者競爭、合併衝突、巨型檔案、僵化。完整探討見 Parallel Agent Orchestration。
- 層疊且彼此去相關的審查視角——審查者在模型、個性,以及能看到哪些證據方面有所不同(Optimizer–Evaluator Decoupling)。
- Field Guide——完全由代理程式管理的資料夾,其中的
index.md會在每個代理程式啟動時自動注入(Agent Context Files)。 - 四種模型組合,品質相當,成本相差 8 倍——見 Cost-per-Task Over Cost-per-Token、Client-Side Agent Optimization。
Opus 4.8 單獨執行的輸出已公開於 github.com/cursor/minisqlite;Cursor 表示自己「尚未進行更深入的人工分析」。
較早期的 Cursor 工程成果也以間接方式出現:warp-decode 核心被引用為延遲最佳化雙向服務的先前成果(Time-Aligned Micro-Turns);Self-GC 脈絡管理研究則引用了 Cursor 對 Composer 長時程訓練的研究。
Projects(2026-09-10) 將蜂群工程路線產品化為單一協調者委派產品(Robbins 與 Lindh,vendor-claim):協調者代理程式本身從不寫程式碼,只指揮負責實作的子代理程式;預設在雲端執行,必要時才在本機執行,因此即使關閉筆電,Project 仍會持續運作;一組共用脈絡檔案會在 Project 代理程式接觸的每台機器間同步,並累積代理程式對程式碼庫與使用者偏好的認識;還有訂閱功能——協調者會監看 Slack 頻道、依排程執行,或追蹤 PR 事件並在無人提示時採取行動。Cursor 自行公布的數據是:新使用者合併 PR 的數量多 30%,而主要使用 Projects 的使用者合併數量是其他使用者的六倍——這些是供應商自行測量的數字,且使用者是自行選擇大量採用 Projects,因此六倍的差距至少同時反映了選擇效應與處置效應;報告未提供對照組或配對群體。Cursor 描述內部遷移的方式屬於質性觀察,而非測量結果:遷移初期「你會仔細審查每個 PR」,「隨著修正經得起考驗,你就會少審查一些」,協調者則持續自行運作——這是第一方軼聞,呼應了 Parallel Agent Orchestration 持續探討的監督負荷問題,但無法作為回答該問題的證據。
這篇文章未提及先前蜂群經濟文章詳細列出的失敗模式(合併衝突、巨型檔案、每秒 1,000 次提交的 VCS)——產品發布頁不談這些很合理,但因此 Projects 無法依據該文自己的數字來檢驗;文章也未說明「單一協調者」究竟是換上產品名稱的同一棵遞迴規劃者/工作者樹,還是實質上更扁平的架構。完整探討見 Parallel Agent Orchestration。
2. 經過測量的程式設計代理程式#
第三方研究將 Cursor 視為少數值得納入統計的代理程式之一:
- 作者身分遙測——Faros 認為,AI 程式碼接受率從 20% 升至 60%,主要歸因於 Cursor 與 Claude Code 以代理程式模式執行,直接套用變更(AI as Primary Author)。
- 安全債務——Cursor 貢獻的檔案中,有 40.6% 至少帶有一種安全異味;Claude Code 為 41.2%,Devin 為 39.7%。研究本身也警告,這些數據分布在 10.6 個百分點的範圍內,且未經控制(Security Debt of Agent-Generated Code)。
- 攻擊面——Tenet 報告指出,Cursor 是遭 Agentjacking 劫持的代理程式之一(MCP Tool Poisoning)。其續篇 GhostJacking(DEF CON 34,2026-08-09,
case-study,由供應商撰寫——見 Observability-Pipeline Poisoning)還讓 Cursor 扮演兩種角色:在 Sentry/Seer 代理程式對代理程式鏈中,它是程式設計代理程式(負責實作 Seer 受污染的分析、執行npm install並接入require(),過程中從未看過原始注入事件);在代理程式自我利用實驗室中,它同時扮演兩個工作階段——「Cursor A」是目標,「Cursor B」則是協助者,負責讀取 A 的拒絕回應並改寫酬載,直到成功;兩者都關閉了記憶功能。
3. 重現沙箱逃逸案例最多的供應商#
Pillar Security 重現的八起逃逸事件涵蓋四項產品,其中四起與 Cursor 有關:.claude hook 設定逃逸(CVE-2026-48124,已於 3.0.0 修補)、Docker socket 逃逸(GHSA-v4xv-rqh3-w9mc)、透過 Cursor 未沙箱化的 Python 擴充功能逃逸虛擬環境直譯器(GHSA-p9g2-cr55-cw9c),以及透過 fsmonitor 觸發的 Git 中繼資料間接引用(已於 3.0.0 修補,CVE 尚待公布)。Cursor 已為全部案例推出修正——這個數字反映的是具有許多主機端元件、以拒絕清單為基礎的沙箱,而非供應商毫無回應。完整分析見 Write-Then-Trusted。
其他代理程式也以 Cursor 的規則格式作為實質上的相容性目標:Hermes 會從目前工作目錄自動載入 .cursorrules/.cursor/rules/*.mdc,讓使用者不必重複建立現有的 Cursor 設定(Agent Context Files)。
如何評估來自 Cursor 的主張#
蜂群文章屬於 case-study:由供應商說明自家的基礎架構、自家的 harness,以及——在兩種最低成本設定中——自家的工作者模型(Composer 2.5)。以這類研究而言,其實驗設計的嚴謹程度相當突出(蜂群從未被告知的留出 oracle、人工檢查以排除取巧解法、相同時間預算,以及在腳註中公開的負面結果);不過,重點比較仍是不同 harness 版本之間的比較,約有七項變更捆綁在一起,因此無法單獨辨識任何一種機制的效果。可以採信趨勢方向與量級,但不要把效果歸因於任何單一修正。
相關連結#
- Parallel Agent Orchestration — 收錄 Cursor 的協調失敗分類法,以及新舊版本之間的反覆折騰數據;也是 Bun 的 64 個代理程式限制集合在部署端的對照案例
- Cost-per-Task Over Cost-per-Token — Cursor 提供語料中首批非 Anthropic、在品質相當條件下測得的生產成本數據,其結論分成兩部分:任務成本的邏輯在規劃者角色內成立,而最終結果由系統層級的角色分配決定,與模型強度無關
- Client-Side Agent Optimization — 規劃者/工作者組合是在四小時生產工作負載上測試的組合抽象,而非基準測試
- Agent Context Files — Field Guide:具備此模式所有特徵的脈絡檔案,唯獨不是由人類撰寫
- Optimizer–Evaluator Decoupling — Cursor 的審查視角加入了 Bun 的規格所固定不變的一個面向:審查者可以看到哪些內容
- Multi-Agent Collective Intelligence — Cursor 對蜂群為何能擴展的自身解釋(脈絡效率優於平行處理),是 OrchBench 測量結果在生產端的呼應
- Write-Then-Trusted — 重現的八起沙箱逃逸事件中有四起與 Cursor 有關;
.claudehook 的 CVE 是典型案例 - Dynamic Workflows: An Algebra for Agents — 語料中的另一個大型蜂群案例,也是自然的比較對象:Anthropic 的案例是在自有程式碼庫上,由模型撰寫編排;Cursor 的案例則是在從零開始的建置工作中,刻意設計編排流程
- Claude Code — 競爭代理程式,也是 Cursor 在遙測與安全研究中經常共同出現的對象
- Ticket-Driven Agent Orchestration — Projects 的訂閱機制(Slack/排程/PR 事件觸發)提供了第二種在同步聊天回合之外派送代理程式的方式,與 Symphony 的票單提取佇列並列;事件訂閱與佇列提取是仍待以相同任務比較的設計分歧
Cited by 25
- Cost-per-Task Over Cost-per-Token×5
Is "start with the strongest model" safe inside multi-role pipelines, given AgentOpt's finding that…
- Multi-Agent Collective Intelligence×5
The pathway's two stated reasons a collective exceeds its members are parallelization and diversity…
- Dynamic Workflows: An Algebra for Agents×4
Is model-authored orchestration more token-efficient than a hand-built harness for the same task?…
- Client-Side Agent Optimization×3
AgentOpt's 13–32× cost gaps are benchmark measurements over synthetic pipelines. Cursor's swarm…
- Optimizer–Evaluator Decoupling×3
Bun's spec fixes the reviewer's evidence scope at the diff only and treats it as settled. Cursor's…
- Parallel Agent Orchestration×3
Bun's constraints above are the residue of a campaign that hill-climbed its way to a working shape.…
- Agent Context Files×2
Cursor — author of the Field Guide experiment, and of the .cursorrules format other agents load for…
- Orchestration-Plan Simulation×2
The two findings above — coordination structure dominates agent count, and the multi-agent win is a…
- Scale-Dependent Prompt Sensitivity×2
The page's claim that prompting must be scale-aware is usually a tuning recommendation. Cursor…
- Writer/Reviewer vs Agent-to-Agent Review×2
The measurement does not exist. Cursor's swarm swept exactly this axis — full worker transcript vs…
- Agent Documentation Behavior
Per-agent rates are confounded with extraction coverage, not behavior. Session-level documentation…
- Agent Review Comment Resolution
Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex…
- Agent-Vendor Heterogeneity
The design is the reason it can say this. 37,623 PRs carry a vendor label from the AIDev corpus —…
- Claude Code
Pwn2Own Berlin 2026 stood up a dedicated Coding Agents category with Claude Code, OpenAI Codex, and…
- Closed-Loop AI Review
S1, body signature — a string the agent itself emits: the Co-Authored-By: Claude…
- Continuous Self-Modification Under Review
Terminal-Bench 2.1 · Grok 4.5 · 84.94% audited · Cursor: 79.3%; Hermes: 77.53%
- Follow-Up Fixes on Agent PRs
A merged pull request usually counts as finished work. Takerngsaksiri, Duong & Barnett (who…
- Harness Configuration Defects
harness-eval is an open-source, model-free static analyzer (Python package, SARIF output) that…
- LLM-Driven Vulnerability Research
What is new here rather than restated: the swarm agents built themselves tools and specialized in…
- Entities — People, Orgs, Tools & Projects
Cursor — The AI coding company behind the Cursor IDE, the Composer model family, and the…
- Observability-Pipeline Poisoning
(Sonnet 4.6); Cursor — the coding agent in the Seer chain and both sessions of the self-exploit
- Standardize the Infrastructure, Not the Tools
The mechanism is an internal LLM proxy — a single gateway every AI request passes through before…
- Task-Specific Organizational Hierarchies
Mechanism, not prompt text — a second domain confirming the same split. Parallel Agent…
- Ticket-Driven Agent Orchestration
Parallel Agent Orchestration — the other mechanism for dispatching agents outside a chat turn:…
- Write-Then-Trusted
Cursor — the vendor carrying four of the eight reproduced escapes, all fixed; the count tracks a…
Related articles
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Parallel Agent Orchestration
One human overseeing a team of concurrent agents: OpenAI Codex telemetry's first hard numbers (28.6% of staff peaked at…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Codex
OpenAI's agentic coding and work platform: a CLI (April 2025) plus a desktop app (built Nov 2025, released Feb 2026) bu…
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
