資料來源#
摘要#
Anthropic 400K 個工作階段研究的縱向發現是:僅僅七個月(2025 年 10 月 → 2026 年 4 月),人們使用 Claude Code 的方式便出現了可量化的轉變——從救火轉向端到端的代理式工作,而且工作的價值也提高了。最明顯的單一變化是:用於修復損壞程式碼的工作階段比例從 33% 降至 19%(除錯比例幾乎減半)。取而代之的是程式碼周邊工作的成長——操作軟體從 14% 增至 21%,而寫作+資料分析大約翻倍,從約 10% 增至約 20%。與此同時,平均工作階段的估計經濟價值上升約 25–27%。這是從使用面觀察到的 harness shrinkage:模型能力提升後,使用者花更少時間處理失效問題,更多時間委派完整任務。
證據註記。
empirical,與 Returns to Expertise in Agentic Coding 相同的一手資料限制(Anthropic 透過 Clio+Sonnet-4.6 分類器測量自家產品,並以遙測驗證;不包含無頭模式/SDK/IDE 的使用情況)。任務價值是刻意設計的相對代理指標——見下文。
九種工作模式#
每個工作階段都會被分類為最能描述它的單一模式。這套分類法本身值得記錄:
- 直接處理程式碼(≈56%): 建立新程式碼(25%)、修復損壞程式碼(26%)、測試/編排(5%)
- 操作(17%): 部署、設定、執行管線、監控
- 規劃/探索(14%): 理解既有系統、規劃變更
- 分析/文字(13%): 分析資料、透過文件/簡報進行溝通
分類器與自動遙測結果一致——它標記為建立/修改程式碼的工作階段中,超過 90% 確實出現了程式碼變更;這是信任其餘以逐字稿為基礎之測量結果的驗證依據。
轉變量化(2025 年 10 月 → 2026 年 4 月)#
| 工作模式 | 方向 | 變化 |
|---|---|---|
| 修復損壞程式碼 | ↓ | 33% → 19% |
| 操作軟體 | ↑ | 14% → 21% |
| 寫作+資料分析 | ↑ | 約 10% → 約 20% |
而且價值全面上升(自由工作市場代理指標):平均工作階段+約 27%;建立**+43%、操作+34%、修復+32%**。報告強調,美元數字較為粗略,最好解讀為相對變化,而非字面上的價值。
其解讀是:反應式、低槓桿工作的時間變少了(除錯),主動式、高槓桿且更自主工作的時間變多了(操作系統、分析資料、撰寫文件)。報告明確將 Claude Code 的使用方式描述為:隨著代理嵌入非程式碼工作,可能是知識工作未來走向的預覽——寫作/分析翻倍,正是這種擴散超越程式碼的前沿。
遙測而非問卷——以及它建立的對照#
如同 Faros,這項研究從系統訊號(此處為工作階段逐字稿+透過 Clio 取得的自動遙測)讀取行為,而非開發者的感受。但兩項遙測研究得出了感覺相反的結論,而這種對照很有啟發性:
- 本研究: 工作階段價值上升約 27%、除錯下降、成功率上升——在個別互動工作階段層級呈現樂觀圖景。
- Faros: 產出上升但品質下降(每位開發者的錯誤+54%、每個 PR 的事件+243%、審查時間 5 倍)——在組織 SDLC 層級呈現悲觀圖景。
這並非直接矛盾;它們測量的是不同單位與階段。Anthropic 測量的是這個工作階段是否達成目標,以及它有多大價值(工作階段內、互動式使用的代理指標,且人在迴圈中);Faros 測量的是數週後整個組織的管線下游發生了什麼(事件、審查佇列、重新開啟的工單)。在工作階段層級為真者,可以同時在組織層級為真——「任務完成了,而且感覺有價值」與「累積的下游成本正在上升」,正是 Telemetry vs. Survey Measurement 所指出的感受生產力與系統結果之分。證據權重也不同:這是 empirical 研究遙測;Faros 則是 vendor-claim 的潛在客戶開發遙測。
相關連結#
- Harness Shrinkage as Models Improve — 除錯比例崩落,以及轉向端到端委派,是從使用資料讀出的 harness shrinkage:模型改善後,每項任務所需的鷹架/救火更少
- Acceleration Whiplash — 並列的遙測結果:此處工作階段層級的價值/成功率上升,對照彼處組織層級的品質下降;單位不同,但兩者都可能為真
- Telemetry vs. Survey Measurement — 兩項研究都測量行為而非感受;這是 Faros
vendor-claim遙測的empirical近親,而工作階段與組織的區分也是相同的感受與系統之分 - Returns to Expertise in Agentic Coding — 同一研究的配套發現:誰能成功,對照本頁探討的什麼正在轉變
- AI as Primary Author — 更多端到端代理式使用(操作/分析/寫作),正是作者身分轉向代理的使用面表現
- Task Time-Horizon Scaling — 能力上限提高是上游原因:更長且可靠的任務時間跨度,讓使用方式得以從修復轉向操作/分析完整工作流程
- Claude Code — 本文追蹤其使用組成的產品
- Conversation-to-Delegation Shift — OpenAI/Codex 跨族群遙測所呈現的相同「詢問→執行」轉變;本頁則是它在 Anthropic/Claude-Code、單一工具脈絡下的姊妹篇——兩個實驗室、兩項工具、同一種轉變
開放問題#
- 觀察窗口只有七個月,而價值代理指標粗略且相對。這 27% 的增幅有多少是真實的任務複雜度成長,又有多少來自分類器/市場配對的漂移?
- 研究排除了無頭模式/SDK/IDE 的使用——這是「相當大的比例」,而且可能是最自動化、最端到端的部分。納入後會加速還是逆轉組成轉變?
- 如果「修復」持續下降,原因是模型較少造成破壞,還是損壞程式碼的工作正轉移到本研究看不到的非互動式管線?
資料來源#
- Agentic coding and persistent returns to expertise — §"What people use Claude Code for"、§"The work"
Cited by 13
- Anthropic Economic Index×2
Agentic coding and persistent returns to expertise (June 2026) — the 400K-session…
- Conversation-to-Delegation Shift×2
The central thesis of OpenAI's The Shift to Agentic AI: Evidence from Codex (Johnston, Holtz,…
- Acceleration Whiplash
Agentic Coding Work Composition Shift — the juxtaposed telemetry: Anthropic's same-period study…
- AI as Primary Author
Agentic Coding Work Composition Shift — more end-to-end agentic use (operate/analyze/write) is the…
- Anthropic
2026-06-16 — Anthropic Economic Research published Agentic coding and persistent returns to…
- Claude Code
Returns To Expertise / Planning Execution Division Of Labor / Agentic Coding Work Composition Shift…
- AI Coding Practice
Agentic Coding Work Composition Shift — Anthropic's 400K-session telemetry, Oct 2025→Apr 2026: as…
- Open Questions Backlog
Agentic Coding Work Composition Shift ×3 (oldest 56d) — The window is seven months and the value…
- Post-Acceptance Edit Behavior
Agentic Coding Work Composition Shift — the usage-composition shift measured one layer down and one…
- Returns to Expertise in Agentic Coding
Agentic Coding Work Composition Shift — the companion finding from the same study: what the work is…
- Task Crossover
Agentic Coding Work Composition Shift — Anthropic's within-tool version: what a session is for…
- Task Time-Horizon Scaling
Agentic Coding Work Composition Shift — the rising reliable-task-length ceiling is the upstream…
- Telemetry vs. Survey Measurement
Agentic Coding Work Composition Shift — the empirical cousin: Anthropic's Clio-based 400K-session…
Related articles
- Planning / Execution Division of Labor
Anthropic's 400K-session telemetry: in a typical Claude Code session humans make ~70% of planning decisions (what to do…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Conversation-to-Delegation Shift
OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production…
- Engineer PM Convergence
Generalists across disciplines; product taste as bottleneck skill; Anthropic Claude Code team as case study; "just do t…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated
Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoreti…
