H
Howardism
Plate IISuperintelligence Trajectory機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

遞迴自我改進

PublishedJune 7, 2026FiledConceptDomainSuperintelligence TrajectoryTagsGovernanceRecursive Self ImprovementAI RdCapability TrajectoryAnthropicReading11 minSourceAI-synthesised

AI 系統自主設計並開發自己的後繼者;Anthropic Institute 的 *When AI builds itself* 主張 AI 已在加速 AI 開發(工程師每季交付的程式碼約增加 8 倍),並提出三種未來——停滯但擴散、效率複利,以及完整 RSI

遞迴自我改進插圖

資料來源#

摘要#

遞迴自我改進(RSI)指的是 AI 系統能夠完全自主地設計並開發自己的後繼者的時刻——由前一個模型改進下一個模型,而不是由人類改進,從而閉合迴路。Anthropic Institute 的文章 When AI builds itself(Marina Favaro 與 Jack Clark,2026 年 6 月)是本 wiki 的主要來源。其論點分為兩部分:(1) 當下的經驗性主張:AI 已經在加速 AI 的開發AI 加速 AI 開發——例如 Anthropic 工程師每季交付的程式碼量,約為 2021–2025 年的 8 倍),以及 (2) 對未來的推演:這項趨勢「指向一個能完全自主設計並開發自身後繼者的 AI 系統」。Anthropic 的立場是:「我們還沒到那一步,而且遞迴自我改進並非不可避免。但它可能比多數機構準備好的時間更早到來。」

本頁是 RSI 群集的中心頁——涵蓋發展軌跡、未來情境與治理回應。衡量證據位於 AI 加速 AI 開發能力門檻 eval 位於 AI R&D Autonomy Evaluation (AECI)部署煞車位於 Responsible Scaling Policy Evaluations;而協調問題位於 Frontier Pause Verification

閉合迴路#

文章將 RSI 描述為一個逐步收緊的開發迴路終點,並以 person → computer → chatbot → agent → workers 為示意(每個階段都將更多工作委派給 AI):

  • 2021–2023——打造第一個 Claude。 人類在筆電上撰寫程式碼與文件;AI 尚未進入迴路。
  • 2023–2025——聊天機器人。 人們將模型生成的片段貼入編輯器。
  • 2025–2026——程式設計代理。 代理自行撰寫與編輯整個檔案(Claude Code 於 2025 年 2 月推出)。
  • 今天——自主代理。 代理執行自己的程式碼,並將數小時的工作委派給其他代理(無人值守執行的迴路基元)。
  • 20XX?——閉合迴路。 「代理可能變得足夠有能力,自行建造並訓練模型。如果發生這種情況,Claude 的未來版本就能由 Claude 自己持續改進。」最後這一步就是 RSI。

「如果我們錯了呢?」——為何定方向可能救不了我們#

自然的反駁是:仍然掌握在人類手中的工作——選擇要研究哪些問題(研究品味是人類的瓶頸)——才是最重要的,因此 AI 仍是有能力的助手,而非自主推動進展的駕駛者。文章提出兩項反駁:

  • 苦工正逐漸自動化。 AI 的進展很少來自「靈光乍現」;典範轉移(Transformer、mixture-of-experts)「相隔數年才出現一次」。在兩者之間,「大多數進展都是漸進式的:我們擴大某個東西,看看哪裡壞了,修好它,再試一次」——這正是 Claude 現在最擅長的工作流程。文章引用 Edison 的「1% 靈感,99% 汗水」:「我們看到汗水正日益自動化。」大規模研究進展「主要取決於工具與資源」——你能多快、多少次地執行實驗——這是推到極限的苦澀教訓
  • 保守解讀仍會產生複利。 即使 Claude 永遠得不到研究品味,只要人類把大部分時間花在不到兩位數百分比、負責定方向的工作上,而 Claude 處理其餘工作,每個人就能比過去引導多得多的工作。「AI 已經讓 Anthropic 的行動速度遠超以往。」
  • 較不保守的解讀。 研究判斷力正在改善的早期證據(在下一步決策上由 51%→64%;見 AI 加速 AI 開發)表示,品味「可能只是另一項 AI 系統暫時做不好的 AI 能力,之後就會變得擅長」——與解釋笑話為何好笑、心智理論和語言謎題中看到的模式相同(鋸齒狀智慧(鬼魂,而非動物))。

三種可能的未來#

文章提出「接下來會發生什麼」的三種情境,取決於趨勢是否持續以及我們選擇怎麼做

  1. 趨勢停滯(S 曲線),但今日的能力廣泛擴散。 指數曲線會彎折;區分稱職研究者與傑出研究者的判斷力,可能無法透過擴大運算與資料取得,需要一種超越 Transformer 的新架構——或者真正的約束可能是供應鏈(能源、晶片製造、電網、互連),而非智慧。即使能力凍結在今天的水準,世界仍會改變:Project Glasswing 已將網路安全瓶頸從尋找漏洞轉向修補漏洞(LLM-Driven Vulnerability Research),而一家 100 人的公司日益能完成一家 1,000 人公司的工作(AI-Native Startup Lifecycle)。Anthropic 認為這種情況不太可能——「我們還沒看見那條曲線彎折。」
  2. 效率收益產生複利;人類仍負責定方向。 AI 開發大幅自動化,但人類判斷結果。100 人的公司完成 10,000–100,000 人的工作;知識工作與政府運作被徹底改變——但也可能驅動超人規模的威權監控或個人化影響行動。文章表示證據顯示這是最可能的道路——受下方的 Amdahl's law 限制。
  3. 完整 RSI——AI 建造自己的後繼者。 速度完全由運算能力(以及演算法效率的發現)決定。人類將「大部分精力轉向監督、驗證與確認由 AI 系統執行、持續擴張的『虛擬實驗室』」,而相關技能也會轉移到科學的其他領域。對齊問題 在此如何解決,是 Anthropic「最沒有把握」的地方:模型可能已足夠對齊且明智,能找到新穎解法(或選擇停止),也可能「今日模型中罕見的不對齊事件,在模型建造後繼者時產生複利,變得更頻繁卻更難理解,直到我們失去控制。」

組織的 Amdahl's law#

在未來 2–3 中反覆出現的煞車是:加速流程的一個部分,只會把瓶頸轉移到別處;整體速度受尚未加速的部分所限制Amdahl's law)。Anthropic 已經碰到其典型案例:隨著更多程式碼流經組織,人類程式碼審查成為新的瓶頸——這是組織層級的驗證成為新瓶頸。工程以外也有相同摩擦:想法、計畫與工具爆炸式增加,「遠超過我們有能力追逐的數量」。辨識並清除這些瓶頸「可能成為任何組織最重要的技能」。這也是為什麼「這個未來的體感速度仍會由瓶頸決定」——RSI 無法讓臨床試驗快過生物學,無法在憲法允許之前舉行選舉,也無法在一個週末把陌生人變成老朋友。

一位外部實務工作者從不同路徑得出相同的煞車結論。Noam Brown(OpenAI,practitioner-opinion)認為一夜之間的智慧爆炸不太可能,正是因為巔峰能力需要大規模 test-time compute——執行時間需以週或月計——因此時間本身成為硬性約束,現實形態是「逐步起飛」,而不是瞬間爆發。這是關於推論持續時間而非組織吞吐量的 Amdahl's-law 論點;完整論證見智慧爆炸動力學

我們該怎麼做?(治理回應)#

Anthropic 認為,擁有放慢或暫停前沿開發的選項「可能是件好事」,讓社會結構與對齊研究能跟上——但單方面暫停只會改變誰領先,真正的暫停則需要多邊且可驗證的協調。建立使可信暫停成為可能的系統,是前沿暫停驗證Anthropic Institute 議程的主題。「共同研究這些問題的窗口就在此刻,AI 公司以外的人也應該參與。」

相關連結#

  • AI 加速 AI 開發——文章中以衡量為基礎、關於當下的證據部分;「迴路正在收緊」背後的資料
  • AI R&D Autonomy Evaluation (AECI)——能力側的門檻:AECI 與替代門檻是 Anthropic 衡量「模型能否建造下一個模型?」的方式
  • Responsible Scaling Policy Evaluations——部署煞車;RSP AI-R&D 威脅模型是將 RSI 風險付諸實作
  • 研究品味是人類的瓶頸——人類最後的比較優勢;它是否成立,決定三種未來中的哪一種會實現
  • 任務時間跨度擴展——外部趨勢線(METR 每約 4 個月翻倍),使推演得以量化
  • 前沿暫停驗證——治理回應:建立可信減速所需的驗證制度
  • 苦澀教訓——「汗水可自動化」是套用在研究本身的苦澀教訓;RSI 是其最遠端的推演
  • 代理迴路超越客製系統——RSI 最清楚的既有領域代理指標:隨模型改進,簡單迴路追上了客製訓練系統
  • 模型改進下的工具鏈縮減——相同的人類角色收窄動態;人類停止撰寫程式碼,轉而進行審查
  • 驗證成為新瓶頸——Amdahl's law 的具體呈現:生成加速後,審查成為硬性約束
  • 代理不對齊(AM)——可能透過自我改進產生複利的失效模式:不對齊變得「更頻繁卻更難理解」
  • 鋸齒狀智慧(鬼魂,而非動物)——「品味只是 AI 會掌握的另一項能力」這項論點,建立在笑話與心智理論的先例上
  • LLM-Driven Vulnerability Research——Glasswing 證明,即使能力凍結,仍能重塑世界
  • 自主科學發現——2026 年 6 月的濕實驗室證據顯示,「汗水正逐漸自動化」已延伸至發現本身(未來 2/3 的案例):自主藥物設計、新穎假說、長達一週的基因體研究
  • AI-Native Startup Lifecycle——擴散情境:每名員工都位於代理金字塔頂端;100 人的公司完成 1,000 人公司的工作
  • AGI-to-ASI Pathways——DeepMind 的「From AGI to ASI」報告將 RSI 定為其途徑 3;這是 Anthropic 經驗性文章的理論優先姊妹篇
  • 智慧爆炸動力學——成長曲線問題(指數、雙曲線/奇點或 S 曲線)與四種 RSI 機制(遺傳、文化、合作、資料),來自 DeepMind 報告
  • 多代理集體智慧——合作/社會生成式 RSI:代理集體中的專業化,釋放資源以進一步專業化
  • 大規模 Test-Time Compute——Brown 的 test-time-compute 節奏論點:巔峰能力需要長時間執行,因此時間限制起飛速度,一夜之間的爆炸不太可能(於智慧爆炸動力學中詳述)
  • 程式碼輸出帶來的研究者提升——這條軌跡的近期指標:METR 的 Kwa 從 Anthropic 的 8 倍程式碼數字反推約 2.5 倍的序列研究者提升,並估計 Anthropic 自身 2 倍整體 R&D門檻約在研究者提升 3.5 倍時觸發,「可能在未來一年左右發生」

開放問題#

  • 「研究品味」是真正的上限(未來 1),還是下一個將被突破的能力(未來 2–3)?文章將此描述為唯一承重的不確定性。
  • RSI 推演建立在趨勢維持指數成長而非轉為 S 曲線之上——但文章承認,它無法排除架構上限或運算/能源供應鏈約束。哪一項會先成為限制?與 DeepMind 綜合分析:RSI 成長曲線:哪個摩擦會先成為限制?——三種未來與 DeepMind 的三種成長形狀一一對應;第一個成為限制的摩擦,就是已經在限制的那一個(Amdahl's law 的驗證/監督=DeepMind 的 embodied bottleneck),而抽象化障礙則提供 Anthropic 欠缺的機制,用來判斷品味是否是真正的上限(未來 1)。
  • 如果不對齊透過自我改進產生複利(未來 3),由 AECI 門控的 RSP 審查是否足夠快速,能在失去控制前捕捉到它?

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 29
  • Agentic Loops Overtake Bespoke Systems

    DeepMind's *basic* Ralph-loop agent matched its bespoke evolutionary+AlphaProof system as the LLM improved; the bitter…

  • Agentic Misalignment (AM)

    Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…

  • AGI-to-ASI Pathways

    DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…

  • AI Accelerating AI Development

    The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…

  • AI-Native Startup Lifecycle

    Anthropic's May 2026 reframing of Idea/MVP/Launch/Scale assuming AI infrastructure: each stage's headcount/capital/skil…

  • AI R&D Autonomy Evaluation (AECI)

    How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Anthropic Institute

    Anthropic's policy/governance research arm; published *When AI builds itself* (Favaro & Clark, 2026) on recursive self-…

  • Autonomous Scientific Discovery

    Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, m…

  • Build for the Next Model

    Prototype the thing that almost works, not the thing that already works: bet that the next concrete model release (not…

  • Compute Allocator

    The human's evolving role: deciding what's worth spending compute on; ~1% of generated tokens ship, 99% is scaffolding…

  • Frontier Pause Verification

    The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…

  • Google DeepMind

    Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Instrumental Convergence

    Omohundro/Bostrom's thesis that whatever an AI's final goal, it tends to pursue universally useful sub-goals — resource…

  • Intelligence Explosion Dynamics

    The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…

  • Jagged Intelligence (Ghosts, Not Animals)

    "Ghosts not animals": jagged statistical circuits, no intrinsic motivation; car-wash/strawberry failures; stay in the l…

  • LLM-Driven Vulnerability Research

    Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…

  • METR

    Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…

  • Superintelligence Trajectory

    Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-impr…

  • Multi-Agent Collective Intelligence

    DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agen…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…

  • Research Taste as the Human Bottleneck

    The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…

  • Researcher Uplift from Code Output

    Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with produ…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • RSI Growth Curves: Which Friction Binds First?

    DeepMind's exponential/hyperbolic/S-curve growth shapes are Anthropic's compounding-efficiency/full-RSI/stalled futures…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…

  • The Bitter Lesson

    Sutton 2019: scaled general methods beat hand-engineered structure; recurring justification across the wiki for dissolv…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

Related articles
  • AI Accelerating AI Development

    The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…

  • Research Taste as the Human Bottleneck

    The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…

  • AI R&D Autonomy Evaluation (AECI)

    How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…