資料來源#
摘要#
When AI builds itself 中提出的治理回應是:如果 RSI 軌跡成立,世界至少應保有減速或暫時暫停前沿 AI 發展的選項,讓社會結構與對齊研究能夠跟上。但只有在暫停具備可信度——是多邊的且可驗證——時,暫停才有用,因為單方面暫停只會改變領先者。 Anthropic Institute 宣稱的議程,是建立可信減速所需的系統。這是 RSP 內部部署煞車的政策對偶:RSP 管控單一實驗室的發布;暫停驗證則是實驗室之間、國家之間的協調問題。
為什麼單方面暫停還不夠#
Anthropic 的立場是:「如果減速只是讓最不謹慎的行動者在技術上追上來,可能會讓所有人都變得更不安全。」單一實驗室單方面暫停「可以立即實現,但作用小得多:它會改變誰是領跑者,卻不會創造目前缺少的更廣泛審議過程。」Anthropic 表示,如果其他前沿或接近前沿的開發者也以可驗證的方式暫停,它會減速或暫時暫停——這使驗證成為關鍵所在。
為什麼 AI 的驗證異常困難#
可信的暫停需要多個資源充足、分布於多個國家的實驗室,在相同條件下同意停止,且各自都能驗證其他實驗室確實停止。AI 甚至讓可探測性(比完整可驗證性更低的門檻)比其他技術更困難:
- 訓練執行比飛彈發射井更容易掩藏。 沒有可供觀察的大型實體特徵。
- 輸入具有通用性。 算力、資料與人才並非武器專用,因此無法像管控裂變材料那樣管控前置要素。
- 暗中背叛的誘因極其龐大——「其他人暫停時,仍繼續的人可能會取得領先地位。」
- 可信的暫停還必須明確規定什麼會觸發暫停、什麼會解除暫停,以及由誰裁決——目前都尚未定義。
先例與時間問題#
這「原則上不一定不可能」——世界曾為複雜技術建立驗證機制,例如 Intermediate-Range Nuclear Forces (INF) Treaty。但這些機制「花了數十年才建立基礎設施與信任」,而在 RSI 的時間表上,「我們沒有那麼久」。因此 Institute 的押注是:現在就開始建立可探測性/驗證基礎設施,早於任何協議,如此在需要時便保有這個選項。未來幾個月內,Anthropic 計畫召集政策制定者、研究人員、公民社會與其他 AI 公司,並發布成果——明確邀請非 AI 公司人士參與審議。
相關連結#
- Recursive Self-Improvement — 讓建立暫停選項值得投入的軌跡;這是它的治理回應
- Responsible Scaling Policy Evaluations — 單一實驗室的部署煞車;暫停驗證是其多邊對應物
- AI Accelerating AI Development — 加速複利的證據,讓「我們沒有數十年」成為實際限制
- Agentic Misalignment (AM) — 失去控制是可信暫停旨在防範的下行風險
- AGI-to-ASI Pathways — DeepMind 的「刻意減速」阻力(阻力 #6)正是同一個協調問題,而其「軍事—經濟適應主義/無政府狀態作為架構師」分析,則說明了可驗證的多邊協調為何如此困難
- Open-Weight Elicitation Irreversibility — 驗證框架中的盲點:暫停可觀察的訓練執行,對已發布權重上的無界推論毫無作用
開放問題#
- AI 訓練的「驗證機制」具體由什麼組成——算力核算、資料中心檢查、硬體認證、晶片遙測?文章點出了問題,卻沒有提出機制。
- 可探測性 < 可驗證性:當訓練執行不留下實體特徵、而輸入具有雙重用途時,探測甚至能否變得可靠?
- 由誰裁決觸發與解除?目前沒有任何機構持有這項授權,而建立這樣的機構本身就是以十年為尺度的任務。
資料來源#
- When AI builds itself — §"What should we do?"(可驗證的多邊暫停;可探測性與可驗證性之別;INF Treaty 先例;Anthropic Institute 的召集活動)
Cited by 10
- AGI-to-ASI Pathways
DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Anthropic Institute
Anthropic's policy/governance research arm; published *When AI builds itself* (Favaro & Clark, 2026) on recursive self-…
- Superintelligence Trajectory
Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-impr…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- RSI Growth Curves: Which Friction Binds First?
DeepMind's exponential/hyperbolic/S-curve growth shapes are Anthropic's compounding-efficiency/full-RSI/stalled futures…
Related articles
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
