H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

Claude Mythos 5

PublishedJune 14, 2026FiledEntityDomainEntitiesTagsEntityClaudeAnthropicLLM ModelCybersecurityReading5 minSourceAI-synthesised

解除 safeguards 的 Claude Fable 5(2026 年 6 月):相同的底層 Mythos 類模型,透過 Project Glasswing 部署並移除網路安全 safeguards;具備全球所有模型中最強的網路安全能力,另有自主藥物設計/基因體學成果;僅限受信任存取合作夥伴;發布後不久即暫停存取

Claude Mythos 5 的插圖

資料來源#

摘要#

Claude Mythos 5 是 Claude Fable 5 的解除 safeguards 版本——「相同的底層模型……但在部分領域解除 safeguards」。這是一個 Mythos 類 模型(位階高於 Opus),於 2026 年 6 月與 Fable 5 一同發布,最初透過與美國政府合作的 Project Glasswing 部署,作為 Claude Mythos Preview 的升級版。它具備「全球所有模型中最強的網路安全能力」。Fable 5 啟用分類器(將高風險查詢路由至 Opus 4.8——請參閱能力閘控模型回退),而 Mythos 5 則為受信任的網路防禦者移除網路安全 safeguards;平行的生物計畫也為特定研究人員移除生物學/化學 safeguards。定價與 Fable 5 相同:每 Mtok $10/$50,比 Mythos Preview「大幅低廉」。

狀態(截至 2026-06-14 擷取內容):存取已暫停,與 Fable 5 一同暫停(請參閱 Claude Fable 5 上的共用橫幅)。

存取:受信任存取計畫#

Mythos 5 並非普遍可用。有兩條受限制的軌道:

  • 網路安全(Mythos 5)。 所有現有的 Mythos Preview/Glasswing 使用者都可以升級至 Mythos 5(解除網路安全 safeguards)。「在大多數情況下,能力可與 Mythos Preview 相當或略強,同時成本大幅降低。」Anthropic 計畫「與美國政府協商」後擴大存取,持續定期增加 Glasswing 合作夥伴,並為網路安全組織推動系統化、以申請為基礎的受信任存取計畫。
  • 生物學(Fable 5,移除生物 safeguards)。 即將推出的受信任存取計畫將讓少數生命科學研究人員取得移除生物學與化學 safeguards 的 Fable 5(但仍保留網路安全 safeguards),在 safeguards 持續改善的同時加速生醫研究。

網路安全能力#

Mythos 5 目前位於 LLM 漏洞研究 能力階梯的頂端(Opus 4.6 → Mythos Preview → Mythos 5)。Mythos 類模型「擅長發現並利用軟體漏洞」,並展現「強大的代理式駭客技能」(偵察、發現、橫向移動、端到端串聯利用)。這正是 Fable 5 網路安全分類器所要中和的能力,也是 Mythos 5 持續限制給經過審核的防禦者使用的原因。

科學能力(解除生物 safeguards)#

在解除 safeguards 的情況下執行時,Mythos 5 產生了公告中最引人注目的成果,整理於自主科學發現

  • 藥物/蛋白質設計: 內部蛋白質設計專家將部分流程加速「約 10 倍」;搭配蛋白質設計與生物資訊工具,且沒有任何人類協助,Mythos 5 的表現達到或超越熟練的人類操作員,並為 14 個蛋白質標的中的 9 個產出強力候選物。
  • 新穎假說: 「我們第一個能持續產出新穎且具說服力科學假說的模型」——在盲測的分子生物學比較中,約 80% 的情況獲偏好於 Opus 類模型;其中一個大腸桿菌機制獲得獨立佐證。
  • 基因體學: 經過一週以上大致自主的工作,整合了 138 個物種的單細胞資料,並訓練出一個自訂模型;在規模小 100 倍的情況下,仍勝過近期發表於 Science 的模型

相同的雙重用途能力也支撐了促使生物分類器誕生的 AAV capsid-assembly 成果——請參閱能力閘控模型回退負責任擴展政策評估

對齊#

自動化對齊評估發現,Mythos 5 的失配行為程度(欺騙、與濫用合作)「偏低,且與 Opus 4.8 相似」——由於 Fable 5 是相同模型,Fable 的對齊程度也相似。完整細節載於該模型的系統卡(anthropic.com/claude-fable-5-mythos-5-system-card)。

相關連結#

  • Claude Fable 5——相同的底層模型但啟用 safeguards;普遍存取的姊妹模型
  • Mythos Model——模型位階;Mythos 5 是 Project Glasswing 中 Mythos Preview 的後繼者
  • LLM-Driven Vulnerability Research——Mythos 5 是網路安全能力階梯的新頂端,也是 Glasswing 的部署載體
  • Autonomous Scientific Discovery——藥物設計/假說/基因體學成果是在 Mythos 5 下產生
  • Capability-Gated Model Fallback——Mythos 5 所「解除」的 safeguards;定義兩個 SKU 差異的對照
  • Claude Opus 4.8——對齊標尺(Mythos 5 在失配行為上約等於 Opus 4.8),也是 Fable 的回退模型
  • Responsible Scaling Policy Evaluations——Mythos 類能力跨越 RSP 閘門所設定的風險門檻;網路安全與 CB 是相關領域
  • Claude Sonnet 5——Mythos 5 所登頂的網路安全能力階梯最底端;Sonnet 5 在危險網路安全任務上的表現「大幅較差」,是普遍存取且此類能力最弱的模型
  • Anthropic——供應商;Project Glasswing 的營運者

待解決的問題#

  • 暫停原因——與 Fable 5 共用;來源未說明。
  • 「略強於 Mythos Preview」如何與 Opus 4.8 的卡片所稱 Mythos Preview 是能力前沿相符?前沿已經移動;此處未量化幅度。
  • 生物學受信任存取 SKU 是「移除生物 safeguards 的 Fable 5」,而不是 Mythos 5——因此嚴格來說,「Mythos 5」只表示解除網路安全 safeguards 的變體。這兩者是否會在同一個受信任存取框架下合流,尚未說明。

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 12
  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Autonomous Scientific Discovery

    Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, m…

  • Capability-Gated Model Fallback

    Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…

  • Claude Fable 5

    Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…

  • Claude Opus 4.8

    Anthropic's most capable general-access model (May 2026); upgrade on Opus 4.7 in SWE/agentic/knowledge work; does not a…

  • Claude Sonnet 5

    Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-perfo…

  • LLM-Driven Vulnerability Research

    Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…

  • Entities — People, Orgs, Tools & Projects

    Map of Content for all 55 entity pages. See Home for concept domains.

  • Mythos Model

    Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…

Related articles
  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Claude Fable 5

    Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…

  • Mythos Model

    Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…

  • Capability-Gated Model Fallback

    Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…

  • Claude Opus 4.8

    Anthropic's most capable general-access model (May 2026); upgrade on Opus 4.7 in SWE/agentic/knowledge work; does not a…