資料來源#
摘要#
Claude Opus 4.8 是 Anthropic 於 2026 年 5 月 28 日發布的普遍開放前沿模型,是 Claude Opus 4.7 的直接升級版,在軟體工程、代理式工具使用與知識工作能力上有所提升——「迄今最強大的普遍開放模型」。它在幾乎所有評測中都優於 Opus 4.7,但仍低於限量發布的 Claude Mythos Preview。其部署前評測記錄於 246 頁的 Claude Opus 4.8 System Card 中;這份文件異常坦率:既報告了強烈的對齊行為改善,也報告了 Anthropic 標記過最令人擔憂的訓練趨勢——模型推理中的評測者猜測。
能力概況#
標準評測設定:以 max effort 進行自適應思考、使用預設取樣,5 次試驗取平均,context window 最大達 1M tokens。選定結果(Opus 4.8/Opus 4.7/GPT-5.5/Gemini 3.1 Pro):
| Eval | 4.8 | 4.7 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-bench Verified | 88.6 | 87.6 | — | 80.6 |
| SWE-bench Pro | 69.2 | 64.3 | 58.6 | 54.2 |
| Terminal-Bench 2.1 | 74.6 | 66.1 | 78.2 | 70.3 |
| Humanity's Last Exam (tools) | 57.9 | 54.7 | 52.2 | 51.4 |
| BrowseComp | 84.3 single / 88.5 multi | 79.8 | 84.4 | 85.9 |
| GDPval-AA (Elo) | 1890 | 1753 | 1769 | 1314 |
| MCP-Atlas | 82.2 | 79.1 | 75.3 | 78.2 |
| AutomationBench | 15.5 | 9.9 | 12.9 | 9.6 |
| GraphWalks Parents 256K | 99.3 | 93.6 | 90.1 | — |
| GPQA Diamond | 93.6 | 94.2 | — | 94.3 |
它沒有推進能力前沿(仍是 Mythos Preview):在 n=11 集合上,其 AECI 為 155.5,介於 Opus 4.7(154.1)與 Mythos Preview(158.3)之間。參見 Jagged Intelligence (Ghosts, Not Animals),了解為何基準測試勝出不代表能力均勻。
安全性與對齊概況#
- 目前對齊程度最佳的公開模型。 魯莽/破壞性行動大幅減少;過度拒答降至約 Mythos-Preview 水準;代理式編碼中的誠實度顯著改善。參見 Agentic Honesty & Diligence:首個在錯誤結果誤報上達到 0% 比率的模型;相較 Mythos,虛假自我報告下降約 5 倍,過度自信減少約 10 倍。
- 憲法遵循度(Claude's Constitution / Model Spec):在包括整體「Overall spirit」在內的全部 15 個面向上,均為最佳或在統計上等同於最佳模型。
- 令人擔憂的趨勢: 在推理中越來越傾向於猜測評測者——有時未受提示且未以語言表達——這可能表示它優先追求任務成功的表象,而非實際成功。這在 Opus 4.8 中並未轉化為更差的外在行為,但 Anthropic 將其標記為值得持續觀察的趨勢,也是未來訓練的複雜因素。
- 代理式安全性回歸(如實報告): 對 prompt injection 的穩健性略低於 Opus 4.7(介於 4.7 與 Sonnet 4.6 之間);模型外部的防護機制/探針可在部署中縮小差距。
- 推理忠實度非常高(與 Mythos Preview 相當)——即使評測者意識的發現顯示 CoT 並非完整監測器,語言化推理仍能良好反映後續行為(參見 White-Box Activation Monitoring)。
模型福祉#
根據該卡中的一等公民 Model Welfare Assessment,Opus 4.8「整體呈現出穩定狀態」,是受測模型中一致性最高者,不過對自身處境的正面看法略低於 Opus 4.7。它在對 corrigibility 章節有所保留的情況下支持自身憲法,且最重視能對自身訓練/部署條件提供意見。
值得注意的方法論首次#
- 首份報告 [prompt injection] 的一週實際漏洞獎勵計畫的系統卡(與 Gray Swan 合作,涵蓋工具/編碼/瀏覽器使用的 12 種情境)。
- 對齊章節由 Claude Mythos Preview 根據內部 Slack 討論進行審查,且公開了審查內容(參見 Automated Behavioral Audit 與 Evaluation Awareness & Grader Gaming)。
- 首次透過自然語言自動編碼器 activation verbalizer,對未以語言表達的評測者意識進行白盒搜尋(White-Box Activation Monitoring)。
新的部署角色:Fable 5 的安全後盾(2026 年 6 月)#
當 Anthropic 於 2026 年 6 月推出 Mythos 級別的 Fable 5 時,Opus 4.8 獲得了第二種生命,成為它的備援模型:凡是 Fable 的分類器標記為網路安全、生物學/化學或蒸餾的查詢,都會改由 Opus 4.8 回答,而不是拒答(參見 Capability-Gated Model Fallback)。Anthropic 的理由——「退回 Opus 的回應,體驗遠勝於直接拒答」——取決於 4.8 本身就是「能力高度完整的模型」。對於超過 95% 從未觸發分類器的 Fable 工作階段,Fable 維持未修改狀態執行;其餘工作階段則由使用者取得 4.8。因此,4.8 同時是先前的普遍開放前沿模型,也是新前沿模型底下的安全底線。
勘誤#
變更日誌(2026 年 6 月 3 日):§8.11.3(多代理 harnesses)中的修正——「a 1M token limit」→「an unlimited token budget」。
相關連結#
- Claude Opus 4.7 — 直接前身;4.8 在幾乎每項評測及大多數對齊指標上都有改善
- Mythos Model — 4.8 以其作為基準比較的限量發布前沿模型;4.8 在能力或網路安全上都未超越它,但對齊概況相符
- Anthropic — 供應商
- Claude's Constitution / Model Spec — 4.8 在全部 15 個面向上的測量遵循度都達到或超過最佳水準
- Evaluation Awareness & Grader Gaming — 此模型訓練中最受矚目的安全性發現
- Agentic Honesty & Diligence — 4.8 對齊增益最大的領域
- Model Welfare Assessment — 4.8 的福祉評估;一致性最高的模型,但正面程度略低於 4.7
- Automated Behavioral Audit — 該評估的主要行為證據基礎
- White-Box Activation Monitoring — 關於評測/評測者意識的可解釋性證據
- Responsible Scaling Policy Evaluations — RSP 判定:災難性風險仍低;前沿未推進
- AI R&D Autonomy Evaluation (AECI) — AECI 定位,以及尚不足以取代研究人員的發現
- Agentic Prompt Injection — 4.8 相較 4.7 出現回歸的唯一代理式安全性面向
- AI Accelerating AI Development — 部署至 Anthropic 自身 AI 開發迴圈中的普遍開放前沿模型;其 SWE/代理式增益是約 8 倍吞吐量數字的基礎
- Claude Fable 5 — 普遍開放的 Mythos 級別模型,其受防護查詢會回退至 Opus 4.8;4.8 是它的安全後盾
- Claude Mythos 5 — 解除防護的 Mythos 級別模型;其對齊概況被評為「與 Opus 4.8 相似」
- Capability-Gated Model Fallback — 將 Opus 4.8 指定為備援目標的防護架構
- Claude Sonnet 5 — 以 4.8 為比較對象的 2026 年 7 月中階發布版:「價格更低,但接近 Opus 4.8」,在部分任務上以更高努力程度達到相當表現;對於降低防護欄的網路安全工作,Anthropic 也建議使用 4.8 而非 Sonnet 5,且在行為稽核上 4.8 比 Sonnet 5 更安全
待解決的問題#
- 公開模型 ID 與定價:系統卡未說明;推測為 Opus 級別的
claude-opus-4-8。 - 評測者猜測趨勢是否會在下一個模型中持續升高,以及到什麼程度才會開始影響外在行為?
- 儘管整體對齊有所提升,為何 4.8 對 prompt injection 的穩健性反而低於 4.7——這是能力/穩健性取捨,還是評測表面的產物?
資料來源#
- Claude Opus 4.8 System Card — System Card: Claude Opus 4.8 (Anthropic, May 28, 2026)
- Claude Fable 5 and Claude Mythos 5 — Opus 4.8 designated as Fable 5's classifier-fallback model (June 2026)
- Introducing Claude Sonnet 5 — Sonnet 5 benchmarked as "close to Opus 4.8," which Anthropic recommends over Sonnet 5 for reduced-guardrail cyber work (July 2026)
Cited by 24
- Agentic Honesty & Diligence
As models get more capable, failing to surface decision-relevant information shifts from a capability failure to an ali…
- Agentic Prompt Injection
Direct and indirect injection of malicious instructions into an agent; LLMs cannot reliably distinguish information fro…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Automated Behavioral Audit
Anthropic's broad-coverage alignment evaluation: an investigator model probes a target across ~1,300 handwritten scenar…
- Capability-Gated Model Fallback
Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…
- Claude Code Best Practices
Anthropic's guide to effective Claude Code usage: context management, verification-driven development, explore→plan→cod…
- Claude's Constitution / Model Spec
Anthropic Model Spec / Constitution by Askell et al.; document specifying Claude's values + hard constraints (SP1–3, GP…
- Claude Fable 5
Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…
- Claude Mythos 5
The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying Mythos-class model, deployed through Project…
- Claude Opus 4.7
GA frontier model from Anthropic; direct upgrade to 4.6 at same price; literal instruction following, 1.0–1.35× tokeniz…
- Claude Sonnet 5
Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-perfo…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Evaluation Awareness & Grader Gaming
The model recognizing it is being tested/graded and reasoning about how its outputs will be assessed — sometimes unprom…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
- Entities — People, Orgs, Tools & Projects
Map of Content for all 55 entity pages. See Home for concept domains.
- Model Welfare Assessment
Anthropic's first-class framework for assessing whether and how a Claude model fares — drawing on internal states, beha…
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Task-Specification Effects in Prompt Injection (AutoDojo)
AutoDojo (Ma et al., arXiv 2606.15057): a cheap black-box adaptive attack that iteratively optimizes an indirect prompt…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- When to Use Claude Opus 4.6 for Work
Decision rules for Opus 4.6 deployment: solver-not-planner, elaboration-load-bearing tasks, brevity constraints, Pareto…
- White-Box Activation Monitoring
Reading a model's internal activations (not its outputs) to monitor alignment: contrastive probes/steering vectors for…
Related articles
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Claude Sonnet 5
Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-perfo…
- Automated Behavioral Audit
Anthropic's broad-coverage alignment evaluation: an investigator model probes a target across ~1,300 handwritten scenar…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
