H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

Mythos 模型

PublishedMay 6, 2026FiledEntityDomainEntitiesTagsEntityModelAnthropicReading8 minSourceAI-synthesised

Anthropic 預覽級前沿模型,也是 Mythos 級別(高於 Opus)的首個成員;出於安全考量受到門控,與 Opus 4.7 一同供內部使用;其後代 Fable 5 / Mythos 5 於 2026 年 6 月推出,成為首批可普遍使用的 Mythos 級別模型

Mythos 模型插圖

資料來源#

摘要#

Anthropic 的預覽級前沿模型。特別的是,它被描述為「極其強大」,並且在安全審查後才開放(Mythos Preview / red.anthropic.com 的發布文章奠定了 LLM-Driven Vulnerability Research 故事的基礎)。目前與 Claude Opus 4.7 一同在 Anthropic 內部使用。截至 2026 年 5 月,尚未 GA。

目前公開所知#

  • Mythos Preview 展現了湧現的網路安全能力——自主發現零日漏洞、完成完整的漏洞利用鏈。詳細分析請參閱 Mythos Preview 發布文章中的 LLM-Driven Vulnerability Research
  • Anthropic 的回應:Project Glasswing 防護措施(在 4.7 發布時被稱為「首批 Glasswing 後防護措施」)。
  • Boris Cherny:「我們會用一點 Mythos 來嘗試,然後大量使用 Opus 4.7 進行內部試用,並撰寫大部分程式碼。」—— Mythos 屬於預覽級,並非主力模型。
  • Cat Wu:「Mythos 是一個極其強大的模型。但我們確實在內部使用這些模型,我認為這讓我們的發布速度有所提升,但我不認為這能解釋增長的大部分。」—— 證實了內部使用,也明確否認它是開發節奏提升的主要解釋。

在 Opus 4.8 System Card 中的角色#

Opus 4.8 System Card(2026 年 5 月)讓 Mythos Preview 的角色變得異常具體——它仍是普遍使用模型用來對照的能力前沿,並且在評估本身中被當作一項工具使用:

  • 前沿基準: Opus 4.8「並未將能力前沿推進到 Mythos Preview 之上」。在 AECI 指數上,Mythos 得分 158.3,高於 Opus 4.8 的 155.5(以及 4.7 的 154.1)。其 Risk Report 界定了 4.8 的 RSP 案例範圍。
  • 調查模型: Mythos Preview 是驅動 Opus 4.8 Automated Behavioral Audit 的兩個調查模型之一(另一個是僅提供協助的 Opus 4.7)。
  • 評估審查者: 在一個值得注意的元層次安排中,Mythos 獲准存取內部 Slack 討論,並受邀審查接近定稿的對齊章節;其(已發布的)審查確認了坦率程度,並指出沒有任何 eval 專門測試訓練作弊——請參閱 Evaluation Awareness & Grader Gaming
  • 對齊標尺: Opus 4.8 在大多數衡量項目上符合 Mythos Preview 的對齊特徵,並在數項誠實度指標上超越它(Agentic Honesty & Diligence)。

能力數據點(When AI builds itself#

Anthropic Institute 的文章(2026 年 6 月)為 Mythos Preview 附上了具體數字;它正是推動 Anthropic AI-R&D 加速(AI Accelerating AI Development)的模型:

  • 時間跨度: METR 評定它能夠工作「至少」16 小時,「已達到 [METR] 在不新增任務的情況下所能測量的上限」。(Task Time-Horizon Scaling)
  • 核心最佳化 eval: 在讓小型模型更快訓練的任務上,2026 年 4 月達到約 52× 加速;相較之下,Opus 4 在一年前約為 3×,人類基準約為 4×——「不到一年內,從超級有幫助進步到超越人類」。
  • 研究下一步判斷: 在困難的繞道路徑時刻中,64% 的時間勝過人類選擇(Opus 4.5 於 2025 年 11 月為 51%)。
  • 自我回報的提升: 在 2026 年 3 月對 130 名研究團隊員工進行的調查中,相較於不使用 AI,受訪者估計使用 Mythos Preview 的產出中位數約為 (Anthropic 認為實際提升幅度略低)。

這些是部署端數據(Mythos 在內部使用),與上述 System Card 對受門控能力前沿的描述有所不同。

為何維持門控#

以下因素的綜合作用:

  • 相較先前模型的網路安全能力差距(依 Mythos Preview 發布文章)
  • 安全機制仍在評估與強化中
  • Anthropic 宣示的使命立場

意味著該模型供內部使用並選擇性預覽,而非廣泛推出。預計 Mythos 的某個後代日後會公開推出——Boris Cherny:「它會以某個版本、某個後代的形式,在某個時間點對所有人開放。」

更新——後代已推出(2026 年 6 月)#

Boris 的預測成真了。2026 年 6 月,Anthropic 推出了 Fable 5Mythos 5——首批可普遍使用的 Mythos 級別模型,也實現了「Mythos 級別」作為位於 Opus 級別之上的具名能力層級。如今的譜系是 Mythos Preview(2026 年 4 月)→ Fable 5 / Mythos 5(2026 年 6 月)

  • Fable 5 = 一個透過分類器「為普遍使用而確保安全」的 Mythos 級別模型;當遇到網路安全、生物或蒸餾查詢時,分類器會回退至 Opus 4.8Capability-Gated Model Fallback)。
  • Mythos 5 = 相同的底層模型,但解除防護措施,透過 Project Glasswing 部署,作為 Mythos Preview 的升級版(「與其相當,或稍強一些」,價格不到一半)。現有的 Glasswing/Mythos-Preview 使用者可直接升級。

這讓能力前沿超越 Mythos Preview——也就是 Opus 4.8 System Card 視為上限的那條線——並且是 Mythos 級別能力首次觸及公眾。(兩者據報在推出後不久都被暫停;請參閱 Claude Fable 5。)

相關連結#

  • Anthropic — 供應商
  • Claude Opus 4.7 — 先前的 GA 模型;Mythos 是下一層級的預覽模型
  • Claude Opus 4.8 — 以 Mythos 為基準的現行 GA 模型;Mythos 是它的能力前沿、行為審查中的調查模型,以及對齊評估的審查者
  • LLM-Driven Vulnerability Research — 能力特徵的主要公開說明
  • Harness Shrinkage as Models Improve — Mythos 級別能力讓 Boris 的「100 行」預測變得可想像
  • AI Native Product Cadence — 明確否認是開發節奏提升的解釋,但確實有所貢獻
  • AI Accelerating AI Development — Anthropic 可量測的 AI-R&D 加速背後的模型(52× 核心 eval、64% 下一步、調查約 4×)
  • Task Time-Horizon Scaling — Mythos 位於時間跨度曲線可測量邊界(16 小時以上)
  • METR — 將 Mythos 的時間跨度評定為「至少 16 小時」,超出其標準測量上限
  • Claude Fable 5 — 首個可普遍使用的 Mythos 級別模型;Mythos Preview 的受防護後代
  • Claude Mythos 5 — 透過 Project Glasswing 部署、解除防護措施的後代,作為 Mythos Preview 的直接升級版
  • Capability-Gated Model Fallback — 讓 Mythos 級別模型得以普遍發布的防護架構
  • Claude Sonnet 5 — Mythos Preview 是自動化行為審查中對齊程度最高的參考,而中階 Sonnet 5 落後於它(Sonnet 5 比 Sonnet 4.6 更安全,但不如 Opus 4.8 與 Mythos Preview)

待解決的問題#

  • Fable 5 / Mythos 5 是否會在推出後暫停後重新上線,以及何時重新上線?
  • 網路安全以外的能力特徵:Mythos Preview 聚焦於安全故事;其他能力面向在外部尚未有充分文件記錄。
  • 內部存取控制:Anthropic 中究竟誰會在日常工作中使用 Mythos,而非 Opus 4.7?Boris 暗示使用頻率不高(嘗試性使用),但未提供詳細資訊。

已解決問題#

  • 公開發布時間表:已回答——Mythos Preview 本身從未 GA,但其後代 Fable 5 / Mythos 5 已於 2026 年 6 月向大眾開放(見上方的後代已推出)。

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 23
  • Anthropic×4

    Mythos Model — first Mythos-class model (Mythos Preview); used internally, gated for safety;…

  • Claude Mythos 5×3

    Claude Mythos 5 is the safeguards-lifted form of Claude Fable 5 — "the same underlying model... but…

  • AI R&D Autonomy Evaluation (AECI)×2

    Mythos Model — the frontier-setting model; its System Card holds the full methodology and bounds…

  • Automated Behavioral Audit×2

    Two: Claude Mythos Preview and a helpful-only variant of Opus 4.7 (expected to be especially good…

  • Claude Fable 5×2

    Claude Fable 5 is Anthropic's first generally-available Mythos-class model (launched June 2026) — a…

  • Claude Opus 4.8×2

    Mythos Model — the limited-release frontier model 4.8 is benchmarked against; 4.8 does not surpass…

  • Claude Sonnet 5×2

    Mythos Model — Mythos Preview is the best-aligned reference on the behavioral audit that Sonnet 5…

  • Cross-Lab Pre-Release Review×2

    Musk's evidentiary anchor is that this has already happened once, informally. Pressed on whether…

  • METR×2

    Long-task measurement at the frontier. METR found Claude Mythos Preview could work for "at least"…

  • Responsible Scaling Policy Evaluations×2

    The Responsible Scaling Policy (RSP) is Anthropic's framework for gating model deployment on…

  • Agentic Honesty & Diligence

    Opus 4.8 is simultaneously (a) the most honest model in outward agentic behavior and (b) the most…

  • AI Native Product Cadence

    Cat is asked directly whether internal access to Mythos explains the velocity:

  • Capability-Gated Model Fallback

    The safeguard architecture that lets Anthropic ship a Mythos-class model for general use: when…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Claude Opus 4.7

    Mythos Model — preview-tier successor used internally; Boris Cherny: "we use a little bit of Mythos…

  • Cost-per-Task Over Cost-per-Token

    Claude Fable 5, Claude Opus 5, Claude Sonnet 5, Mythos Model — the classes being chosen between

  • Evaluation Awareness & Grader Gaming

    As an extra assurance, Anthropic had Claude Mythos Preview review the near-final alignment section…

  • LLM-Driven Vulnerability Research

    Mythos Model — entity page for the preview model that produced these findings; internal use at…

  • Entities — People, Orgs, Tools & Projects

    Mythos Model — Anthropic preview-tier frontier model and the first member of the Mythos-class tier…

  • Motivated Mislabeling

    This is also the closest published thing to the eval gap Mythos Preview flagged when reviewing the…

  • Open Questions Backlog

    Mythos Model ×3 (oldest 98d) — Do Fable 5 / Mythos 5 return after the post-launch suspension, and…

  • Researcher Uplift from Code Output

    Anthropic's Mythos Preview system card said overall R&D acceleration is "well short of a sustained,…

  • Task Time-Horizon Scaling

    Mythos Preview is at the edge of measurability: METR found it could work for "at least" 16 hours…

Related articles
  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Claude Opus 4.8

    Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallba…

  • Claude Opus 5

    Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…

  • Claude Sonnet 5

    Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-perfo…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…