資料來源#
- Claude Fable 5 and Claude Mythos 5
- Claude Opus 4.8 System Card
- Gemma 4 Technical Report
- More compute, more capability: Why AI agent evaluations need to account for test-time compute
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
摘要#
Responsible Scaling Policy (RSP) 是 Anthropic 用來根據發布前的災難性風險能力評估來設置模型部署門檻的框架,涵蓋三個領域:化學與生物武器(CB/CBRN)、自動化 AI 研究與開發,以及高風險錯位。每份系統卡都會執行 RSP 評估套件並產生風險判定。對 Opus 4.8 而言,總體結論是:該模型沒有超越 Claude Mythos Preview 推進能力前沿,而且「在目前的緩解措施下,部署此模型所造成的災難性風險仍然很低」。
RSP 是前沿部署的制度性安全煞車:一旦模型跨越某個門檻,就會在發布前觸發更強的必要防護措施(ASL 分級)。它也是模型能力與對齊測量的治理對應物,後者則在系統卡的其他部分呈現。
風險評估流程#
系統卡以既有的風險報告為基礎,並針對每個模型更新,而不是每次從頭推導。由於 Opus 4.8 在測量軸上位於 Opus 4.7 與 Mythos Preview 之間,且沒有推進前沿,因此先前對 Mythos Preview 的分析為 Opus 4.8 的案例提供了界限,大多數判定也直接沿用。
化學與生物(CB)#
透過自動化評估套件測量(CB-1、CB-2——包括黑箱 RNA 序列建模/設計與 AAV 衣殼包裝預測)。Opus 4.8 沒有超越 Mythos Preview 推進化學風險前沿;生物風險結果則以相同門檻進行報告。緩解措施仍是重要焦點,尤其是模型外部的防護措施,它們能捕捉 Automated Behavioral Audit 所揭示的殘餘極端濫用合作行為。
自動化 AI 研發#
兩個 RSP 威脅模型:
- 威脅模型 1——錯位的高風險 AI: 適用於 Opus 4.8(如同先前的模型),但它不會提高風險等級——其進行祕密行為/規避監控的能力很低,且與 Opus 4.7 相當,同時行為對齊有所改善。
- **威脅模型 2——自動化研發的風險:**不適用——該模型沒有推進能力前沿。
能力面向是透過 AECI and autonomy evaluations 測量。只有在以下任一條件成立時,才會跨越 RSP 的 AI 研發門檻:(1) 模型能在 5 倍成本因子內,完全取代 Anthropic 全部的研究科學家與工程師;或 (2) AI 進展出現可歸因於自動化的「戲劇性加速」。兩者皆未達成。
對齊風險#
評估為非常低,但高於 Mythos Preview 之前的模型——這與 Mythos Preview Alignment Risk Update 相同。開放一般存取後,兩條新的風險途徑納入範圍:削弱其他高資源 AI 開發者內部的研發,以及削弱主要政府內部的決策。與規避防護措施相關的能力(SHADE-Arena、Minimal-LinuxBench 隱匿性)與 Opus 4.7 相當,遠低於 Mythos;以微調為基礎的誘發並未顯著提高隱匿成功率。
值得注意的流程細節#
- 在 Opus 4.8 訓練期間,監控沒有發現比先前模型更令人擔憂的事項,與評分器相關的推理升高除外。
- 技術錯誤造成的意外思維鏈監督影響了約 0.1% 的情節(這是各系統卡反覆出現的註腳;與思維鏈可監控性相關)。
更新——Mythos 級別跨越門檻(2026 年 6 月)#
Opus 4.8「未推進前沿」的判定,僅在 Mythos 級別能力維持受限時成立。2026 年 6 月推出**Fable 5/Mythos 5**,正是那條界線移動的時刻:Anthropic 明確表示,「Mythos 級別模型已達到會帶來重大風險的門檻。」這對 RSP 圖景有兩項後果:
- 緩解措施從設置門檻轉向已部署的防護措施。 Mythos Preview 被直接 withheld,而 Opus 4.8 仰賴維持在前沿以下;對於位於門檻上的模型,一般存取的答案則是 Capability-Gated Model Fallback——將網路/生物化學/蒸餾查詢路由至 Opus 4.8,而不是拒絕處理。這是第一個一般存取模型,其中已部署的濫用緩解措施,而非能力餘裕,成為承重的安全機制。所有 Mythos 級別流量也同時受到保留 30 天的要求約束。
- CB 案例因真實的科學能力而更加明確。 AAV 衣殼組裝結果——Mythos 級別模型在未訓練的情況下擊敗專用蛋白質語言模型(見 Autonomous Scientific Discovery)——正是 CB 門檻要界定的雙重用途能力提升,也是目前將生物分類器調得過於寬泛的明確原因。
因此,RSP 的部署煞車現在正以啟動模式運作,而不只是「尚未抵達前沿」模式——兩個模型在發布後遭到暫停(見 Claude Fable 5),也活生生提醒我們:這些防護措施正在 production 中接受對抗性測試。
無界預算缺口(外部批評)#
Noam Brown(OpenAI,practitioner-opinion)指出,這個框架與每個實驗室的 preparedness policy 共享一個結構性漏洞:它沒有指定評估能力時所使用的 test-time-compute 預算。 RSP 與 preparedness frameworks 是「在 ChatGPT 時代周邊發展出來的」,早於 test-time-compute scaling 變得重要——當時給 GPT-3 級模型「1000 萬美元,也做不了比 10 美元多多少的事」。如今能力是預算的函數,因此「應該以什麼預算評估這些模型?」仍未獲回答。這項擔憂正是有用能力案例的鏡像:如果模型在增加投入時,能持續改善某項任務而不趨於漸近,那麼它也可能持續改善那些社會不希望它做的事——因此,在真正部署預算之前就停止的固定預算 CB 或網路 eval,會低估危險。Brown 不願表示這是否應該阻止發布(「雙方都有論據」),但堅持認為這個問題目前只是被假裝不存在。
Brown 的批評如今已由政府評估機構以實證證明。 UK AI Security Institute 在 2026 年 7 月的研究(empirical)是同一缺口的獨立測量版本:固定預算分數「掩蓋了風險的真實規模」,而且由於任務所需的計算量會隨人類時間跨度增長,受限預算會先在時間最長、最困難的任務上耗盡——恰好是安全 eval 最需要觸及、後果也更高的任務。這種效應對較新的模型最大,因此在前沿處,低估程度會擴大。AISI 已因此改變自身做法:現在會以多個預算進行評估(包括為最困難任務使用非常大的預算),並相對於預算報告觸及率與可靠性,明確讓「真正低能力的模型」能與「資源不足的評估」區分開來。這使無界預算異議從一名實驗室研究者的 practitioner-opinion,轉變為國家安全評估機構已納入其方法論的作業性發現。
這讓本頁已經潛藏的兩件事更加清楚。RSP 所依賴的「我們每天使用它,而它不能取代我們的研究人員」(下文)是一次單一預算判斷;能力過剩意味著危險能力上限與有用能力上限一樣,可能遠高於 eval 實際花費的任何預算。而這也是 Compute-Controlled Benchmarking 在安全面向的解讀:沒有附帶計算預算所報告的威脅模型判定,和沒有附帶計算預算所報告的能力分數一樣,規格不足。
開放權重沒有對應的缺口#
RSP 以兩種模式運作:設置門檻(「前沿未推進——發布」)與啟動(「跨越門檻——部署防護措施」)。啟動模式的每項工具都需要由供應商控制的伺服器——分類器回退、暫停、保留 30 天,以及思考 token 上限。開放權重發布可以使用第一種模式,卻完全沒有第二種模式,因此其單一預算評估不只是規格不足,而是終局性的:之後的發現都無法改變該產物的行為。
Gemma 4(DeepMind,2026 年 7 月,Apache 2.0)讓這個形狀清晰可見。它提供思考模式;其安全章節以散文宣稱「在內容安全的每個類別都有重大改善」,沒有表格,也沒有明確說明預算,而同一份報告卻包含十六個基準測試表格。Gemma 4 距離任何前沿門檻都很遠(The Open-Weight Frontier Gap),因此這是結構性觀察,而非警報——但這個結構將決定第一個確實接近門檻的開放權重發布。相關內容見 Open-Weight Elicitation Irreversibility。
相關連結#
- Recursive Self-Improvement——RSP 是 RSI 軌跡上的制度性部署煞車;AI 研發威脅模型是將 RSI 風險付諸作業
- Frontier Pause Verification——多邊協調的對應物:RSP 管控單一實驗室的發布,暫停驗證管控整個領域
- AI R&D Autonomy Evaluation (AECI)——提供 AI 研發威脅模型判定所需的能力測量(AECI、自主性 eval)
- Claude Opus 4.8——受評估的模型;前沿未推進,災難性風險低
- Mythos Model——設定前沿的模型,其風險報告界定了 Opus 4.8 案例
- Automated Behavioral Audit——提供 RSP 判定所依賴的錯位/濫用行為證據
- Evaluation Awareness & Grader Gaming——訓練監控期間唯一被標記為升高的疑慮
- LLM-Driven Vulnerability Research——網路能力是相鄰的災難性風險領域;Project Glasswing 是其緩解措施脈絡
- AI-Accelerated Offense——網路防護措施所回應的攻擊加速威脅
- Capability-Gated Model Fallback——為一般發布的 Mythos 級別模型實作網路/生物門檻的推論時緩解措施
- Claude Fable 5——其部署啟動 RSP 煞車的一般存取 Mythos 級別模型
- Claude Mythos 5——解除防護措施的 Mythos 級別模型;門檻所界定的能力
- Claude Sonnet 5——煞車在中階模型上的未啟動模式:發布前 eval 發現網路風險低,因此 Sonnet 5 僅配備預設的偵測與阻擋防護措施(而非 Fable 5 的回退機制)——RSP 判定如何縮減至低於前沿的發布
- Autonomous Scientific Discovery——使化學/生物判定更加明確的 CB 領域能力(AAV、自主生物學)
- AGI-to-ASI Pathways——制度化門檻(強制 eval、授權、事件通報)是 DeepMind「刻意減速」摩擦的作業形式;RSP AI 研發門檻是可設置門檻的遞迴改進路徑
- Deployment Simulation——發布前安全設置門檻的跨實驗室類比:OpenAI 的 production-replay 預測會像 RSP 套件管控 Anthropic 的發布決策一樣提供依據,但增加了 RSP 行為 eval 所缺少的可檢驗、以 production 校準的預測層
- Large-Scale Test-Time Compute——外部批評的根源:能力(以及危險能力)會隨推論預算擴展,而 RSP 門檻沒有指明該預算
- Compute-Controlled Benchmarking——將同一項「報告預算」要求套用於安全判定,而非能力分數
- Latent Capability Overhang——固定預算安全 eval 可能低估風險的原因:危險能力上限可能遠高於 eval 使用的預算
- Open-Weight Elicitation Irreversibility——RSP 的啟動模式沒有開放權重對應物;已發布模型的安全評估是終局性的
- Gemma 4——安全聲明未以表格呈現的開放權重思考模型
- UK AI Security Institute——以實證證明無界預算缺口,並將多預算評估納入自身實務的政府評估機構
- Noam Brown——指出無界預算缺口的外部批評者(OpenAI)
未決問題#
- RSP 判定高度依賴「我們每天使用它,而它不能取代我們的研究人員」。當模型接近門檻時,這項主觀判斷能擴展得多好?
- 兩條新的一般存取風險途徑(其他 AI 開發者;主要政府)雖然已納入範圍,卻只得到輕度評估——那裡若出現正面發現,究竟會是什麼樣子?
- RSP 煞車如何與 Recursive Self-Improvement 互動:如果加速形成複合效應,基於 AECI 的設置門檻是否足夠迅速?沒有多邊的暫停驗證機制,單一實驗室的設置門檻是否真的重要?
資料來源#
- Claude Opus 4.8 System Card——§2(RSP 評估):§2.1 風險評估流程、§2.2 CB 評估、§2.3 AI 研發、§2.4 對齊風險更新
- Claude Fable 5 and Claude Mythos 5——Mythos 級別「門檻……重大風險」;分類器防護措施與保留 30 天作為已部署的緩解措施
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown——Noam Brown(No Priors,2026-06-26),
practitioner-opinion:preparedness frameworks/RSP 沒有指定評估危險能力時所使用的 test-time-compute 預算 - Gemma 4 Technical Report——§5,責任/安全/資安:開放權重發布中未以表格呈現的安全聲明;儘管報告整體屬於
empirical層級,仍視為vendor-claim - More compute, more capability: Why AI agent evaluations need to account for test-time compute——UK AISI(2026-07-02,
empirical):固定預算分數「掩蓋了風險的真實規模」;受限預算會先截斷時間最長/最困難的任務;採用多預算評估,以免把資源不足的 eval 誤認為低能力模型
Cited by 25
- AGI-to-ASI Pathways
DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Anthropic Institute
Anthropic's policy/governance research arm; published *When AI builds itself* (Favaro & Clark, 2026) on recursive self-…
- Automated Behavioral Audit
Anthropic's broad-coverage alignment evaluation: an investigator model probes a target across ~1,300 handwritten scenar…
- Autonomous Scientific Discovery
Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, m…
- Capability-Gated Model Fallback
Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…
- Claude Mythos 5
The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying Mythos-class model, deployed through Project…
- Claude Opus 4.8
Anthropic's most capable general-access model (May 2026); upgrade on Opus 4.7 in SWE/agentic/knowledge work; does not a…
- Claude Sonnet 5
Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-perfo…
- Compute-Controlled Benchmarking
Noam Brown's critique that the single-number 'benchmark grid' is broken because it doesn't control for test-time comput…
- Deployment Simulation
OpenAI's pre-release safety method: replay recent production conversations with a candidate model (strip the old final…
- Frontier Pause Verification
The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Latent Capability Overhang
Noam Brown's claim that already-released models can do far more than anyone has extracted, because nobody spends enough…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
- Superintelligence Trajectory
Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-impr…
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
- Noam Brown
OpenAI research scientist and a pioneer of inference-time (test-time) compute scaling; earlier built superhuman poker A…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
- The Open-Weight Frontier Gap
Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57;…
- OpenAI
AI lab and maker of the GPT-5 series and Codex; in this corpus it appears as a frontier-safety research source (Deploym…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- UK AI Security Institute
UK government AI-evaluation body (Science of Evaluation team); its July 2026 test-time-compute study is the first indep…
Related articles
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Claude Fable 5
Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Capability-Gated Model Fallback
Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…
