資料來源#
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
- When AI builds itself
摘要#
遞迴自我改進(RSI)指的是 AI 系統能夠完全自主地設計並開發自己的後繼者的時刻——由前一個模型改進下一個模型,而不是由人類改進,從而閉合迴路。Anthropic Institute 的文章 When AI builds itself(Marina Favaro 與 Jack Clark,2026 年 6 月)是本 wiki 的主要來源。其論點分為兩部分:(1) 當下的經驗性主張:AI 已經在加速 AI 的開發(AI 加速 AI 開發——例如 Anthropic 工程師每季交付的程式碼量,約為 2021–2025 年的 8 倍),以及 (2) 對未來的推演:這項趨勢「指向一個能完全自主設計並開發自身後繼者的 AI 系統」。Anthropic 的立場是:「我們還沒到那一步,而且遞迴自我改進並非不可避免。但它可能比多數機構準備好的時間更早到來。」
本頁是 RSI 群集的中心頁——涵蓋發展軌跡、未來情境與治理回應。衡量證據位於 AI 加速 AI 開發;能力門檻 eval 位於 AI R&D Autonomy Evaluation (AECI);部署煞車位於 Responsible Scaling Policy Evaluations;而協調問題位於 Frontier Pause Verification。
閉合迴路#
文章將 RSI 描述為一個逐步收緊的開發迴路終點,並以 person → computer → chatbot → agent → workers 為示意(每個階段都將更多工作委派給 AI):
- 2021–2023——打造第一個 Claude。 人類在筆電上撰寫程式碼與文件;AI 尚未進入迴路。
- 2023–2025——聊天機器人。 人們將模型生成的片段貼入編輯器。
- 2025–2026——程式設計代理。 代理自行撰寫與編輯整個檔案(Claude Code 於 2025 年 2 月推出)。
- 今天——自主代理。 代理執行自己的程式碼,並將數小時的工作委派給其他代理(無人值守執行的迴路基元)。
- 20XX?——閉合迴路。 「代理可能變得足夠有能力,自行建造並訓練模型。如果發生這種情況,Claude 的未來版本就能由 Claude 自己持續改進。」最後這一步就是 RSI。
「如果我們錯了呢?」——為何定方向可能救不了我們#
自然的反駁是:仍然掌握在人類手中的工作——選擇要研究哪些問題(研究品味是人類的瓶頸)——才是最重要的,因此 AI 仍是有能力的助手,而非自主推動進展的駕駛者。文章提出兩項反駁:
- 苦工正逐漸自動化。 AI 的進展很少來自「靈光乍現」;典範轉移(Transformer、mixture-of-experts)「相隔數年才出現一次」。在兩者之間,「大多數進展都是漸進式的:我們擴大某個東西,看看哪裡壞了,修好它,再試一次」——這正是 Claude 現在最擅長的工作流程。文章引用 Edison 的「1% 靈感,99% 汗水」:「我們看到汗水正日益自動化。」大規模研究進展「主要取決於工具與資源」——你能多快、多少次地執行實驗——這是推到極限的苦澀教訓。
- 保守解讀仍會產生複利。 即使 Claude 永遠得不到研究品味,只要人類把大部分時間花在不到兩位數百分比、負責定方向的工作上,而 Claude 處理其餘工作,每個人就能比過去引導多得多的工作。「AI 已經讓 Anthropic 的行動速度遠超以往。」
- 較不保守的解讀。 研究判斷力正在改善的早期證據(在下一步決策上由 51%→64%;見 AI 加速 AI 開發)表示,品味「可能只是另一項 AI 系統暫時做不好的 AI 能力,之後就會變得擅長」——與解釋笑話為何好笑、心智理論和語言謎題中看到的模式相同(鋸齒狀智慧(鬼魂,而非動物))。
三種可能的未來#
文章提出「接下來會發生什麼」的三種情境,取決於趨勢是否持續以及我們選擇怎麼做:
- 趨勢停滯(S 曲線),但今日的能力廣泛擴散。 指數曲線會彎折;區分稱職研究者與傑出研究者的判斷力,可能無法透過擴大運算與資料取得,需要一種超越 Transformer 的新架構——或者真正的約束可能是供應鏈(能源、晶片製造、電網、互連),而非智慧。即使能力凍結在今天的水準,世界仍會改變:Project Glasswing 已將網路安全瓶頸從尋找漏洞轉向修補漏洞(LLM-Driven Vulnerability Research),而一家 100 人的公司日益能完成一家 1,000 人公司的工作(AI-Native Startup Lifecycle)。Anthropic 認為這種情況不太可能——「我們還沒看見那條曲線彎折。」
- 效率收益產生複利;人類仍負責定方向。 AI 開發大幅自動化,但人類判斷結果。100 人的公司完成 10,000–100,000 人的工作;知識工作與政府運作被徹底改變——但也可能驅動超人規模的威權監控或個人化影響行動。文章表示證據顯示這是最可能的道路——受下方的 Amdahl's law 限制。
- 完整 RSI——AI 建造自己的後繼者。 速度完全由運算能力(以及演算法效率的發現)決定。人類將「大部分精力轉向監督、驗證與確認由 AI 系統執行、持續擴張的『虛擬實驗室』」,而相關技能也會轉移到科學的其他領域。對齊問題 在此如何解決,是 Anthropic「最沒有把握」的地方:模型可能已足夠對齊且明智,能找到新穎解法(或選擇停止),也可能「今日模型中罕見的不對齊事件,在模型建造後繼者時產生複利,變得更頻繁卻更難理解,直到我們失去控制。」
組織的 Amdahl's law#
在未來 2–3 中反覆出現的煞車是:加速流程的一個部分,只會把瓶頸轉移到別處;整體速度受尚未加速的部分所限制(Amdahl's law)。Anthropic 已經碰到其典型案例:隨著更多程式碼流經組織,人類程式碼審查成為新的瓶頸——這是組織層級的驗證成為新瓶頸。工程以外也有相同摩擦:想法、計畫與工具爆炸式增加,「遠超過我們有能力追逐的數量」。辨識並清除這些瓶頸「可能成為任何組織最重要的技能」。這也是為什麼「這個未來的體感速度仍會由瓶頸決定」——RSI 無法讓臨床試驗快過生物學,無法在憲法允許之前舉行選舉,也無法在一個週末把陌生人變成老朋友。
一位外部實務工作者從不同路徑得出相同的煞車結論。Noam Brown(OpenAI,practitioner-opinion)認為一夜之間的智慧爆炸不太可能,正是因為巔峰能力需要大規模 test-time compute——執行時間需以週或月計——因此時間本身成為硬性約束,現實形態是「逐步起飛」,而不是瞬間爆發。這是關於推論持續時間而非組織吞吐量的 Amdahl's-law 論點;完整論證見智慧爆炸動力學。
我們該怎麼做?(治理回應)#
Anthropic 認為,擁有放慢或暫停前沿開發的選項「可能是件好事」,讓社會結構與對齊研究能跟上——但單方面暫停只會改變誰領先,真正的暫停則需要多邊且可驗證的協調。建立使可信暫停成為可能的系統,是前沿暫停驗證與 Anthropic Institute 議程的主題。「共同研究這些問題的窗口就在此刻,AI 公司以外的人也應該參與。」
相關連結#
- AI 加速 AI 開發——文章中以衡量為基礎、關於當下的證據部分;「迴路正在收緊」背後的資料
- AI R&D Autonomy Evaluation (AECI)——能力側的門檻:AECI 與替代門檻是 Anthropic 衡量「模型能否建造下一個模型?」的方式
- Responsible Scaling Policy Evaluations——部署煞車;RSP AI-R&D 威脅模型是將 RSI 風險付諸實作
- 研究品味是人類的瓶頸——人類最後的比較優勢;它是否成立,決定三種未來中的哪一種會實現
- 任務時間跨度擴展——外部趨勢線(METR 每約 4 個月翻倍),使推演得以量化
- 前沿暫停驗證——治理回應:建立可信減速所需的驗證制度
- 苦澀教訓——「汗水可自動化」是套用在研究本身的苦澀教訓;RSI 是其最遠端的推演
- 代理迴路超越客製系統——RSI 最清楚的既有領域代理指標:隨模型改進,簡單迴路追上了客製訓練系統
- 模型改進下的工具鏈縮減——相同的人類角色收窄動態;人類停止撰寫程式碼,轉而進行審查
- 驗證成為新瓶頸——Amdahl's law 的具體呈現:生成加速後,審查成為硬性約束
- 代理不對齊(AM)——可能透過自我改進產生複利的失效模式:不對齊變得「更頻繁卻更難理解」
- 鋸齒狀智慧(鬼魂,而非動物)——「品味只是 AI 會掌握的另一項能力」這項論點,建立在笑話與心智理論的先例上
- LLM-Driven Vulnerability Research——Glasswing 證明,即使能力凍結,仍能重塑世界
- 自主科學發現——2026 年 6 月的濕實驗室證據顯示,「汗水正逐漸自動化」已延伸至發現本身(未來 2/3 的案例):自主藥物設計、新穎假說、長達一週的基因體研究
- AI-Native Startup Lifecycle——擴散情境:每名員工都位於代理金字塔頂端;100 人的公司完成 1,000 人公司的工作
- AGI-to-ASI Pathways——DeepMind 的「From AGI to ASI」報告將 RSI 定為其途徑 3;這是 Anthropic 經驗性文章的理論優先姊妹篇
- 智慧爆炸動力學——成長曲線問題(指數、雙曲線/奇點或 S 曲線)與四種 RSI 機制(遺傳、文化、合作、資料),來自 DeepMind 報告
- 多代理集體智慧——合作/社會生成式 RSI:代理集體中的專業化,釋放資源以進一步專業化
- 大規模 Test-Time Compute——Brown 的 test-time-compute 節奏論點:巔峰能力需要長時間執行,因此時間限制起飛速度,一夜之間的爆炸不太可能(於智慧爆炸動力學中詳述)
- 程式碼輸出帶來的研究者提升——這條軌跡的近期指標:METR 的 Kwa 從 Anthropic 的 8 倍程式碼數字反推約 2.5 倍的序列研究者提升,並估計 Anthropic 自身 2 倍整體 R&D門檻約在研究者提升 3.5 倍時觸發,「可能在未來一年左右發生」
開放問題#
- 「研究品味」是真正的上限(未來 1),還是下一個將被突破的能力(未來 2–3)?文章將此描述為唯一承重的不確定性。
- RSI 推演建立在趨勢維持指數成長而非轉為 S 曲線之上——但文章承認,它無法排除架構上限或運算/能源供應鏈約束。哪一項會先成為限制?與 DeepMind 綜合分析:RSI 成長曲線:哪個摩擦會先成為限制?——三種未來與 DeepMind 的三種成長形狀一一對應;第一個成為限制的摩擦,就是已經在限制的那一個(Amdahl's law 的驗證/監督=DeepMind 的 embodied bottleneck),而抽象化障礙則提供 Anthropic 欠缺的機制,用來判斷品味是否是真正的上限(未來 1)。
- 如果不對齊透過自我改進產生複利(未來 3),由 AECI 門控的 RSP 審查是否足夠快速,能在失去控制前捕捉到它?
資料來源#
- When AI builds itself——Anthropic Institute,When AI builds itself: Our progress toward recursive self-improvement, and its implications(Marina Favaro 與 Jack Clark,2026 年 6 月)
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown——Noam Brown(No Priors,2026-06-26),
practitioner-opinion:由於 test-time-compute 依賴使時間成為硬性約束,一夜之間的爆炸不太可能(「逐步起飛」)
Cited by 29
- Agentic Loops Overtake Bespoke Systems
DeepMind's *basic* Ralph-loop agent matched its bespoke evolutionary+AlphaProof system as the LLM improved; the bitter…
- Agentic Misalignment (AM)
Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…
- AGI-to-ASI Pathways
DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
- AI-Native Startup Lifecycle
Anthropic's May 2026 reframing of Idea/MVP/Launch/Scale assuming AI infrastructure: each stage's headcount/capital/skil…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Anthropic Institute
Anthropic's policy/governance research arm; published *When AI builds itself* (Favaro & Clark, 2026) on recursive self-…
- Autonomous Scientific Discovery
Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, m…
- Build for the Next Model
Prototype the thing that almost works, not the thing that already works: bet that the next concrete model release (not…
- Compute Allocator
The human's evolving role: deciding what's worth spending compute on; ~1% of generated tokens ship, 99% is scaffolding…
- Frontier Pause Verification
The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…
- Google DeepMind
Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Instrumental Convergence
Omohundro/Bostrom's thesis that whatever an AI's final goal, it tends to pursue universally useful sub-goals — resource…
- Intelligence Explosion Dynamics
The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…
- Jagged Intelligence (Ghosts, Not Animals)
"Ghosts not animals": jagged statistical circuits, no intrinsic motivation; car-wash/strawberry failures; stay in the l…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
- METR
Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…
- Superintelligence Trajectory
Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-impr…
- Multi-Agent Collective Intelligence
DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agen…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Research Taste as the Human Bottleneck
The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…
- Researcher Uplift from Code Output
Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with produ…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- RSI Growth Curves: Which Friction Binds First?
DeepMind's exponential/hyperbolic/S-curve growth shapes are Anthropic's compounding-efficiency/full-RSI/stalled futures…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- The Bitter Lesson
Sutton 2019: scaled general methods beat hand-engineered structure; recurring justification across the wiki for dissolv…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
Related articles
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
- Research Taste as the Human Bottleneck
The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
