H
Howardism
Plate IISuperintelligence Trajectory機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

自主科學探索

PublishedJune 14, 2026FiledConceptDomainSuperintelligence TrajectoryTagsGovernanceAI RdScientific DiscoveryCapability TrajectoryDual UseAnthropicReading7 minSourceAI-synthesised

Mythos 類模型如今能在有限人類輸入下進行新穎科學研究——自主蛋白質/藥物設計速度約快 10 倍,達到熟練人類水準;分子生物學假說相較 Opus 類模型約有 80% 更受偏好(其中一項 E. coli 機制獲得獨立佐證);並以小 100 倍的規模,透過長達一週的基因體學研究擊敗一個發表於 Science 的模型;這是 AI 驅動形式化證明搜尋的濕實驗室對應版本,也是研究品味辯論的新證據

自主科學探索插圖

資料來源#

摘要#

透過 Mythos 5(解除生物安全防護的 Fable 5 版本),Anthropic 報告了首批 Claude 結果:模型能夠幾乎獨立地進行新穎科學研究——選擇實驗步驟、執行領域工具、從失敗中恢復,並產出達到或超越熟練人類與近期已發表基準的發現。這是AI 驅動形式化證明搜尋濕實驗室/生命科學對應版本:形式化證明搜尋有 Lean 編譯器作為即時驗證器,而科學的驗證器是實驗——速度較慢且成本更高——因此這裡的主張是經驗性展示與精選範例,而非經編譯器檢查的保證。這些結果是迄今支持遞迴式自我改進較不保守解讀的最有力證據:也就是「苦功正逐漸自動化」已延伸至發現本身,而研究品味可能只是「AI 暫時做不好的另一項能力,之後便會變得擅長」。

三項結果#

藥物/蛋白質設計——達到人類水準的自主性#

Anthropic 的內部蛋白質設計專家使用 Mythos 5,讓藥物設計的部分工作「加速約 10 倍」。在一項研究中,Mythos 5 配備蛋白質設計與生物資訊工具,卻沒有任何人類協助,仍達到或超越熟練人類操作者,執行「科學家通常完成的所有任務:選擇結合位點、選取並執行蛋白質設計工具,以及在過程中從失敗中恢復」。14 個蛋白質目標中有 9 個產生了強效藥物設計候選物,目前正在研究中(免疫檢查點、生長因子/受體訊號傳遞、神經退化、肌肉疾病,以及更棘手的結構性目標)。

新穎假說——優於 Opus 類模型,其中一項獲得佐證#

Mythos 5 是 Anthropic 「首個能持續產出新穎且具說服力之科學假說的模型」。在與 Opus 類模型進行的盲測正面比較中,Anthropic 科學家約 80% 的時間偏好 Mythos 的分子生物學假說,並將其中數項推進至實驗評估。其中一項 Mythos 假說——針對 E. coli 蛋白質提出的新穎機制——獲得了由研究同一問題之實驗室所進行研究的獨立佐證

基因體學——自主工作一週,擊敗規模小 100 倍的已發表模型#

在「超過一週幾乎自主的工作」中,Mythos 5 整理了來自138 種動物、數百萬個細胞的單細胞資料,接著設計並訓練自訂機器學習模型,以辨識即使在親緣關係遙遠的生物中,執行相同角色的細胞。在僅有高層次人類輸入的情況下,該訓練完成的模型**勝過近期發表於 Science 的模型——儘管其規模小了 100 倍。**Anthropic 計畫發表相關成果。

雙重用途的陰影#

正是這項能力,讓一般存取的 Fable 5 必須受到生物學防護。促成這項評估的問題是:預測基因修改如何影響腺相關病毒(AAV)的衣殼組裝——這是基因療法中的真實元件,而其設計能力「落入錯誤的人手中,可能使危險病毒的設計成為可能」。Mythos 類模型在這項任務上勝過專用的蛋白質語言模型,儘管它們並未針對該任務受訓,而是單靠生物學推理。自主科學能力與生物提升風險,是從兩面看待同一項能力——這正是 RSP CB 判定與生物分類器存在的核心張力。

為何這對發展軌跡很重要#

  • 苦功自動化延伸至發現。當 AI 建造自己主張,多數研究進展都是「擴大規模、看看哪裡壞掉、修好它」這類 Claude 擅長的漸進式工作。自主基因體學——整理資料、設計模型、訓練模型、擊敗基準——就是在科學領域端到端執行的這個迴圈,而不只是工程。
  • 它開始侵蝕品味護城河。「持續產出新穎且具說服力的假說」與「僅有高層次人類輸入」,正是原本預期會持續由人類掌握的定方向功能。約 80% 的盲測偏好是一道具體裂縫——但仍由人類評判,且資料來自內部。
  • **仍然崎嶇,仍受驗證所限制。**這些是經過整理的展示(崎嶇智慧(幽靈,而非動物));科學的驗證器是緩慢的濕實驗室確認,而非編譯器,因此不同於AI 驅動形式化證明搜尋,結果無法自動驗證——仍等待實驗與同儕審查。這使它鄰近、但仍低於 Anthropic 所設定門檻的 AI 研發自主性

相關連結#

  • AI 驅動形式化證明搜尋 — 形式數學上的同源案例:AI 進行新穎研究,但擁有即時編譯器驗證器;科學以(緩慢且昂貴的)實驗取代它,因此驗證在此成為更難的瓶頸
  • 遞迴式自我改進 — 支持「苦功正逐漸自動化」這一文章較不保守解讀的最明確濕實驗室證據
  • 研究品味作為人類瓶頸 — 自主假說生成與「僅有高層次人類輸入」,正直接侵蝕人類剩餘的比較優勢
  • AI 研發自主性評估(AECI) — 鄰近的自主性:模型設計並訓練另一個模型、擊敗已發表基準,具有 AI 研發的形狀,但領域是基因體學而非 AI 本身
  • 任務時間範圍擴展 — 「超過一週幾乎自主的工作」是具體的長時間跨度資料點,超出 Mythos Preview 已測得的 16 小時
  • 崎嶇智慧(幽靈,而非動物) — 注意事項:這些是仍然崎嶇之能力的精選展示,不是均勻一致的能力
  • 可驗證性論題 — 限制案例:科學的可驗證性低於 Lean 證明,因此自主性超越了廉價驗證——獎勵訊號是實驗,而非編譯器
  • 能力門控的模型回退 — 雙重用途的另一面;AAV 結果是生物分類器的促成範例
  • 負責任擴展政策評估 — 這些能力所推進的 CB(化學/生物)風險領域
  • Claude Mythos 5 — 產生這些結果的模型(解除生物安全防護)
  • Claude Fable 5 — 生物學受到防護的一般存取同源模型
  • 抽象屏障 — DeepMind 屏障的實際測試:這些結果是否跨越它(新穎原語),或是在由人類定義的空間內運作,而具身瓶頸仍限制著濕實驗室驗證?
  • 變革性創造力 — 自主假說生成是否正從 Boden 式探索性創造力,攀升至變革性(新概念空間)創造力

開放問題#

  • 每項結果都由 Anthropic 報告並挑選範例;基因體學「規模小 100 倍卻擊敗 Science」的主張只是「計畫發表」——外部同儕審查後還能保留多少?
  • 科學的驗證落差:形式證明迴圈能自我驗證;在此,一個錯誤卻充滿自信的假說,必須耗費一輪濕實驗室週期才能證偽。沒有快速驗證器的自主性,是否會增加而非減輕驗證瓶頸?
  • 如果假說生成確實獲得約 80% 的偏好,那麼「研究品味」還剩多少可稱為人類獨有的功能——又該如何測量這項殘餘?

資料來源#

  • Claude Fable 5 and Claude Mythos 5 — §"Evaluating Claude Fable 5 and Claude Mythos 5"(藥物設計;新穎假說;基因體學)與 §"Biology and chemistry"(AAV 雙重用途)
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 16
  • The Abstraction Barrier

    Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives…

  • AI-Driven Formal Proof Search

    LLM generates Lean, compiler verifies every step → eliminates hallucination; DeepMind resolves 9/353 Erdős + 44/492 OEI…

  • AI R&D Autonomy Evaluation (AECI)

    How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Capability-Gated Model Fallback

    Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to…

  • Claude Fable 5

    Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…

  • Claude Mythos 5

    The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying Mythos-class model, deployed through Project…

  • Jagged Intelligence (Ghosts, Not Animals)

    "Ghosts not animals": jagged statistical circuits, no intrinsic motivation; car-wash/strawberry failures; stay in the l…

  • Superintelligence Trajectory

    Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-impr…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…

  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Research Taste as the Human Bottleneck

    The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • RSI Growth Curves: Which Friction Binds First?

    DeepMind's exponential/hyperbolic/S-curve growth shapes are Anthropic's compounding-efficiency/full-RSI/stalled futures…

  • Task Time-Horizon Scaling

    METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…

  • Transformative Creativity

    Boden's three-level model of creativity (combinational, exploratory, transformative) used to locate today's AI achievem…

Related articles
  • Claude Opus 4.8

    Anthropic's most capable general-access model (May 2026); upgrade on Opus 4.7 in SWE/agentic/knowledge work; does not a…

  • Mythos Model

    Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…

  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Intelligence Explosion Dynamics

    The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…