資料來源#
摘要#
Claude Sonnet 5 是 Anthropic「迄今最具代理能力的 Sonnet」(於 2026 年 7 月 2 日宣布),是 Sonnet 4.6 的直接升級版,能制定計畫、使用工具(瀏覽器、終端機),並以「就在幾個月前仍需要更大型、更昂貴模型的水準」自主執行。Anthropic 的定位是:Sonnet 系列開啟了代理時代(3.5 / 3.6 / 3.7 是首批具備強大程式設計與工具使用能力的模型),但近期的代理能力進展集中在 Opus 系列——Sonnet 5 以較低價格縮小了這項差距,表現接近 Opus 4.8。 API 模型 ID:claude-sonnet-5。這是第一方發布公告,因此屬於 vendor-claim 來源;以下基準測試差異均來自 Anthropic,更完整的評估則收錄於 Claude Sonnet 5 System Card。
定價與身分#
- **介紹期定價(截至 2026 年 8 月 31 日):**輸入 token 每百萬 2 美元,輸出 token 每百萬 10 美元。
- **標準定價(自 2026 年 9 月 1 日起):**輸入每百萬 3 美元、輸出每百萬 15 美元;Opus 4.8 則為 5 / 25 美元。
- API 模型 ID
claude-sonnet-5;這是 Free 與 Pro 方案的預設模型,也可供 Max、Team、Enterprise 使用,並可在 Claude Code 與 Claude Platform 上使用。 - Chat、Cowork、Claude Code 與 Claude Platform 全面提高速率限制,以容納更高努力程度的 token 使用量。
能力概況#
標題式主張是成本效能範圍,而非單一峰值。Anthropic 以 Sonnet 4.6(前代模型)與 Opus 4.8(能力更強的參考模型)為基準,在兩項代理 eval 上比較 Sonnet 5——BrowseComp(代理搜尋)與 OSWorld-Verified(電腦使用):
- 在推理、工具使用、程式設計與知識工作方面,全面優於 Sonnet 4.6。
- 成本效能範圍比 Opus 4.8 更寬廣,可透過 effort 參數調整(最高至
xhigh):在中等努力程度下,成本效率顯著更佳;而更高努力程度的執行「在某些任務上可以匹敵 Opus 4.8」。其訴求是,使用者可以在 Sonnet 5 與 Opus 4.8 之間調整努力程度,找出合適的成本/效能平衡——這是 Client-Side Agent Optimization 中努力程度/預算槓桿的第一方實例。 - 早期存取合作夥伴表示,它「能完成先前 Sonnet 模型做到一半便會停止的複雜任務」,並且**「不需明確要求就會檢查自己的輸出」**——這種自發性自我驗證能力,正是 shrinking-harness 論點所預測的現象(原本由 harness 提供的驗證腳手架,會移轉到模型中)。
(Sonnet 4.6 與 Opus 4.8 的正面基準測試表僅以圖片形式發布於來源中,因此此處不轉錄。)
基準測試勘誤(方法論,而非模型):一份 2026 年 6 月 30 日的變更紀錄,將 BrowseComp 成本效能圖表修正為標準代理搜尋方法(含壓縮與程式化工具呼叫的 1,000 萬 token 預算),先前的方法低估了 Sonnet 5。Sonnet 4.6 的基線也重新陳述:Humanity's Last Exam 重新評分為 34.6%(無工具)/46.8%(有工具),OSWorld-Verified 則為 78.5%,原因是 eval 方法變更——因此數字與 Sonnet 4.6 發布部落格中的數字不同。
Token 經濟學(遷移風險)#
Sonnet 5 使用更新後的 tokenizer——這類變更與 Opus 4.7 引入的相同——因此同一份輸入依內容類型而定,會對應到約 1.0–1.35 倍的 token。Anthropic 設定介紹期定價,是為了讓從 Sonnet 4.6 遷移的成本「大致持平」。如同 Opus 4.7,實際流量的 token 膨脹取決於內容,值得實際測量而非假設(可交叉參考 Claude Code Best Practices 中「上下文視窗是首要限制」的主題)。
安全與對齊概況#
Anthropic 的部署前評估顯示,Sonnet 5 整體優於 Sonnet 4.6:
- 代理安全:更擅長拒絕惡意要求,並且能抵抗提示注入攻擊中的劫持嘗試。
- **誠實性:**幻覺與諂媚率低於 Sonnet 4.6。
- 自動化行為稽核(涵蓋多種情境中的合作濫用、欺騙及其他不對齊行為):Sonnet 5 的整體分數低於——也就是更安全於——Sonnet 4.6,但**高於(更差於)能力更強的 Opus 4.8 與 Claude Mythos Preview。**這與通常的擔憂正好相反:在此稽核中,能力更強的模型反而更對齊,而 Sonnet 5 殘留的不對齊是中階能力的產物,不是前沿能力的產物。
網路安全能力與防護措施#
Sonnet 5 並未刻意接受網路安全任務訓練(可對比 Opus 4.7,其網路安全能力在訓練期間被差異化削弱)。它能處理例行且無害的網路安全工作,但在漏洞利用開發等危險任務上,表現「顯著較差」於 Opus 4.8 與 Mythos 5。
- Firefox 漏洞利用 eval(與 Mozilla 共同建立;Firefox 148 中的所有漏洞均已修補):兩個 Sonnet 模型在開發可運作漏洞利用方面的得分都是 0.0%。Sonnet 5 的部分成功率略高於 Sonnet 4.6——Anthropic 將此歸因於通用智慧的提升,而非網路安全專項訓練。
- 預設啟用防護措施。由於 Sonnet 5 在此方面略強於 4.6,因此發布時採用與 Opus 4.7 和 4.8 相同的即時網路安全防護措施(在推論時偵測並阻擋遭禁止/高風險的網路安全使用)。這些防護被判定為低風險,不如 Fable 5 啟用時的防護措施嚴格(後者會阻擋更廣泛的網路安全任務,並切換至較弱的模型)。合法的安全研究人員可透過Cyber Verification Program進行;Anthropic 建議需要降低防護限制的網路安全工作使用 Opus 4.8。
這使 Sonnet 5 成為 Capability-Gated Model Fallback 所描繪防護光譜上的獨特位置:沒有刻意的訓練降級(其低網路安全能力是原生的)、採用 Opus 4.7/4.8 嚴格程度的推論時偵測與阻擋,且沒有模型切換回退——比 Fable 5 的分類器加回退機制更窄,因為其底層能力提升風險被判定為低。
可用性#
自發布日起(2026 年 7 月 2 日)全面可用:Free/Pro 預設使用,Max/Team/Enterprise 可用,可在 Claude Code 與 Claude Platform 上使用(原生、AWS、Microsoft Foundry;Google Vertex 對 Cyber Verification Program 的支援即將推出)。API ID claude-sonnet-5。
相關連結#
- Claude Opus 4.8——Sonnet 5 用來衡量的能力上限:「以較低價格接近 Opus 4.8」、在較高努力程度下於部分任務匹敵它,且在行為稽核上比 Sonnet 5 更安全;也是 Anthropic 建議用於低防護限制網路安全工作的模型
- Claude Opus 4.7——兩項與遷移相關變更的先例:1.0–1.35 倍 tokenizer 膨脹,以及 Sonnet 5 繼承的預設即時網路安全防護
- Claude Fable 5——防護光譜中更嚴格的一端;Sonnet 5 的網路安全防護明確「不如 Fable 5 啟用時的防護嚴格」
- Claude Mythos 5——Sonnet 5 遠遠不及的網路安全能力參考點
- Mythos Model——Mythos Preview 是行為稽核中對齊程度最佳的參考,Sonnet 5 落後於它
- Anthropic——供應商
- Claude Code——主要代理執行環境;Sonnet 5 在發布時即作為可用模型提供
- Capability-Gated Model Fallback——Sonnet 5 為防護光譜框架新增低風險、無回退的一點
- Client-Side Agent Optimization——Sonnet 5 的努力程度成本效能調校,是模型分工/預算/路由槓桿的第一方實例
- Harness Shrinkage as Models Improve——「不需要求就檢查自己的輸出」代表驗證腳手架從 harness 移轉到模型中
- Agentic Prompt Injection——改進的劫持抵抗能力是代理安全的主要提升
- Automated Behavioral Audit——Sonnet 5 接受評分的對齊評估(比 4.6 更安全,比 Opus 4.8/Mythos Preview 更差)
- LLM-Driven Vulnerability Research——Sonnet 5 刻意較弱的網路安全能力軸;Firefox 漏洞利用 eval 是實例
- Responsible Scaling Policy Evaluations——Sonnet 5 的部署前安全/能力 eval,以及其低風險網路安全判定
待解決的問題#
- 來源中的 Sonnet 4.6 與 Opus 4.8 正面基準測試數字僅以圖片呈現;System Card 收錄完整資料。
- 一般 Sonnet 5 流量的實際 token 膨脹倍數是多少(1.0–1.35 倍取決於內容)?當努力程度提高後,「大致持平成本」是否仍然成立?
- 為何中階模型在行為稽核中的不對齊程度高於能力更強的 Opus 4.8 與 Mythos Preview——這是能力與對齊的耦合,還是 Sonnet 與 Opus/Mythos 系列之間訓練配方的差異?
- Sonnet 5 究竟在哪個努力程度上能匹敵 Opus 4.8?其交叉點成本與直接執行 Opus 4.8 相比如何?
資料來源#
- Introducing Claude Sonnet 5 — Anthropic, "Introducing Claude Sonnet 5"(2026 年 7 月 2 日;變更紀錄於 2026 年 6 月 30 日編輯)。
evidence: vendor-claim
Cited by 20
- Cost-per-Task Over Cost-per-Token×4
The one quantified result: on SWE-bench Pro, Sonnet 5 with a Fable 5 advisor lands within 10% of…
- Anthropic×3
Claude Sonnet 5 — mid-tier July 2026 release; most agentic Sonnet yet, default model for Free/Pro…
- Claude Fable 5×2
The same guidance positions Fable as the advisor in the cheap-worker/strong-advisor pattern: Sonnet…
- Claude Opus 4.8×2
Claude Sonnet 5 — the July 2026 mid-tier release measured against 4.8: "close to Opus 4.8 at lower…
- Open Questions Backlog×2
Claude Sonnet 5 ×3 (oldest 41d) — The head-to-head benchmark numbers vs Sonnet 4.6 and Opus 4.8 are…
- Agentic Prompt Injection
Claude Sonnet 5 — improved hijack-resistance is a headline agentic-safety gain over Sonnet 4.6; the…
- Automated Behavioral Audit
Claude Sonnet 5 — scored on the same audit: safer overall than Sonnet 4.6 but worse than the more…
- Capability-Gated Model Fallback
Claude Sonnet 5 — a lower-risk point on the same safeguard spectrum: native low cyber capability…
- Claude Code
Claude Sonnet 5 — available model in Claude Code from launch (July 2026); the cheaper agentic…
- Claude Mythos 5
Claude Sonnet 5 — the far bottom of the cyber-capability ladder Mythos 5 tops: Sonnet 5 performs…
- Claude Opus 4.7
Claude Sonnet 5 — inherits two of 4.7's migration-relevant changes: the 1.0–1.35× tokenizer…
- Claude Opus 5
Claude Sonnet 5 — remains more robust than Opus 5 on raw browser-use injection without safeguards
- Client-Side Agent Optimization
Claude Sonnet 5 — a vendor-shipped instance of the same lever: Anthropic pitches dialing the effort…
- GLM (Z.AI)
The corpus's first third-party coding placement for GLM-5.2, and it is an economic one. Databricks'…
- Instruction Compounding
Design: a block of N ∈ {10, 20, 40, 80, 120, 160} simultaneous, programmatically verifiable rules…
- LLM-Driven Vulnerability Research
Claude Sonnet 5 — the low-capability end of the ladder: 0.0% working-exploit rate on the Firefox…
- Entities — People, Orgs, Tools & Projects
Claude Sonnet 5 — Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at…
- Mythos Model
Claude Sonnet 5 — Mythos Preview is the best-aligned reference on the automated behavioral audit…
- Responsible Scaling Policy Evaluations
Claude Sonnet 5 — the brake's disengaged mode on a mid-tier model: pre-deployment evals found low…
- When to Use Claude Opus 4.6 for Work
> Anthropic's general-access frontier), alongside the Claude 5 family. The
Related articles
- Claude Opus 4.8
Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallba…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Claude Opus 5
Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
- Capability-Gated Model Fallback
Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…
