H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

AlphaProof Nexus

DeepMind 用於 LLM 輔助 Lean 證明生成的框架;包含四種代理程式 (從基本到完整功能);採用 proof-sketch + EVOLVE-BLOCK 介面;SafeVerify。其 44/492 OEIS 結果在 2026-08 成為基準門檻,當時 Epoch AI 在相同項目集上以三工具 ReAct 迴圈、每題 $50 上限解出 147 題;較新兩到三個月的模型是混淆因素

Article metadata
Publication details
Published:May 23, 2026
Filed:Entity
Domain:Entities
Tags:EntitySystemGoogle DeepmindAI For Mathematics
Reading:8 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

AlphaProof Nexus 的插圖

資料來源#

摘要#

Google DeepMind 用於在 Lean 中進行 LLM 輔助形式證明生成的框架(arXiv 2605.22763)。代理程式會查詢前沿 LLM(Gemini 3.1 Pro)與 Lean 編譯器,把 proof sketch(證明部分以 sorry 標記的定理)轉化為經驗證且不含 sorry 的證明。這套框架涵蓋四種代理程式設計,從最簡單的迴圈到演化系統;它也是針對開放研究問題進行首度大規模評估的工具,成果包括 9/353 個 Erdős 問題、44/492 個 OEIS 猜想,以及最佳化、代數幾何、加法組合學、圖論與量子光學領域的結果。Lean 證明已在 github.com/google-deepmind/alphaproof-nexus-results 開源。

四種代理程式(A → D)#

代理程式設計備註
(A) 基本型獨立的證明子代理程式,沒有共享狀態;每個都是由多個回合組成的「Ralph 迴圈」出乎意料地在全部 9 個 Erdős 解題上與 (D) 不相上下——參見 Agentic Loops Overtake Bespoke Systems
(B) 基本型 + AlphaProof(A) 加上可呼叫為工具的 AlphaProof RL 證明器在較難問題上比 (A) 更有效率
(C) 基本型 + 演化(A) 加上族群/Elo 演化式搜尋—
(D) 完整功能型同時採用演化機制與 AlphaProof用於探索開放問題;Evolutionary Proof Search 詳述其運作機制

證明子代理程式會進行多回合 Gemini 3.1 Pro 對話,並使用 search_replace 工具;每次編輯後,Lean 編譯器會回傳回饋,引導下一回合;每個回合結束時,草稿都會由 SafeVerify 驗證(可編譯、不含 sorryAx/公理注入),如果還有 sorry,代理程式就會撰寫經驗教訓註解並繼續。$N$ 個子代理程式平行執行;找到證明時便全部停止。

輸入/輸出:證明草稿#

輸入是一個 Lean 檔案——目標定理以 sorry 標記,並附上所需定義與匯入——也可以選擇加入以 Lean 編碼的自然語言背景與領域知識。可編輯區域由 EVOLVE-BLOCK(輔助引理/定義/證明步驟)與 EVOLVE-VALUE(參數運算式)界定。EVOLVE-VALUE 機制促成了一項真正的發現:將某個最佳化演算法的學習排程標記為值,讓代理程式 D 能同步搜尋排程與證明,進而找到具備更強收斂保證的新參數選擇。

AlphaProof(工具)#

AlphaProof 是 DeepMind 一套獨立且較早推出的系統——以強化學習訓練的奧林匹亞等級 Lean 定理證明器(也是 DeepMind IMO 成果背後的系統)。在 Nexus 中,它被用作專注式證明工具:收到子目標後,會回傳證明、反證或失敗結果。它具備 Test-Time RL 模式,但在此處以成本較低的低運算量樹狀搜尋執行(約 400 次模擬、27.5 TPU 小時 ≈ 每題 $60)。值得注意的是,在獨立樹狀搜尋模式下,AlphaProof 未能解出任何評估問題——它的價值是在 LLM 驅動的迴圈內協助完成子目標,而非單獨作戰。

模型與成本#

多回合證明使用 Gemini 3.1 Pro;評分代理程式則使用較便宜的 Gemini 3.0 Flash。每題推論成本為「數百美元」,且差異很大;使用較小模型(Gemini 3.0 Flash、3.1 Flash-Lite)的基本型代理程式版本一題也沒解出——能力明顯受模型規模限制(見 Scale-Dependent Prompt Sensitivity)。報告的成本不包括 AlphaProof 約 $60 的費用,以及在全部 353 題中尋找可處理問題所耗費的可觀成本。

重要性#

這套系統以編譯器為 LLM 的數學推理提供根據,將容易出現幻覺的自然語言證明轉化為可檢查的產物(AI-Driven Formal Proof Search),並展現隨著 LLM 進步,簡單的代理程式迴圈愈來愈能與特製訓練系統匹敵(Agentic Loops Overtake Bespoke Systems)。已解出的 Erdős 問題記錄在 Terence Tao 的 AI 貢獻維基上。

第三方重新執行 OEIS 結果(2026-08)#

上述 44/492 的 OEIS 數字,如今成了他人圖表中的基準門檻。 OEIS OPEN(OEIS Open: How many conjectures can language models turn into theorems?、Epoch AI, arXiv 2608.11941,empirical)重新整理了這套系統形式化的相同 492 個猜想,建立一套開放且防作弊的評估,任何通用 LM 都能接受測試;他們使用配備三種工具(bash、文字編輯器、預算回報器)的 ReAct 迴圈,並設定每個猜想 $50 的上限,以 Claude Opus 4.8 解出492 題中的 147 題(30%)——是本系統 9% 的 3.3 倍,每個解出的猜想平均花費 $10,與本系統自行估計的每題約 $10 相同;該估計是作者透過私人通信提供給 Epoch 的(最難的少數題目最高約 $50)。

解讀這個數字時須留意三項限制。模型相差一個世代:這些證明子代理程式使用 Gemini 3.1 Pro(2026 年 2 月 19 日),而比較對象是 Opus 4.8 與 GPT-5.5(5 月 28 日與 4 月 23 日);Epoch 自己的註腳也指出了這項差距,因此這項比較所提供的證據,是關於模型發行版本之間的 harness 縮減,而非此架構在固定模型下毫無價值。產物數量與論文不符:Epoch 指出,論文雖報告 44/492,但隨附的 google-deepmind/alphaproof-nexus-results 儲存庫只發布了 38 份 OEIS 證明。Epoch 自己的驗證在一方面更嚴格、另一方面較寬鬆——它使用本系統採用的相同 SafeVerify,但又以 Comparator 交叉核對,因此標題數字從 147 下修為 144。

相關連結#

  • Many-Agent Proof Harnesses — 嘗試不用核心來完成相同工作:Google 的 Stellar Colosseum 協調數十個模型實例進行自然語言證明,以一群對抗式反證者和全域驗證器取代 Nexus 的 Lean 閘門
  • OEIS Open Benchmark — 將本系統自己的 492 題 OEIS 項目集轉化為由第三方建立、開放且防作弊的基準測試;三工具 ReAct 迴圈以相近的單題解題成本解出 147 題,本系統則解出 44 題(受兩到三個月的模型差距影響),並指出實際發布的產物數量是 38,而非 44
  • FrontierMath Erdős Benchmark — 這套系統的 Erdős 探索工作之後,出現了採計分制且預算固定的後續基準:精選 68 道重要問題,每題 $300、72 小時、只能嘗試一次,五種模型中表現最佳者解出 2 題。其解題率(2.9%)與本系統在較難題目集上的 9/353(2.5%)相當,部分原因是 Epoch 的測試規範也將全部項目的搜尋成本計入,而「每題數百美元」並未計入這筆成本
  • Lean — 這套系統使用的證明助理/驗證器
  • Google DeepMind — 背後的研究實驗室
  • Evolutionary Proof Search — 完整功能代理程式 (D) 的族群/Elo 機制
  • Agentic Loops Overtake Bespoke Systems — 本系統自己的基本型代理程式 (A) 在多數問題上與完整系統不相上下
  • Agent Loop Pattern — 基本型證明子代理程式是「Ralph 迴圈」(huntley2025ralph)
  • Client-Side Agent Optimization — A/B/C/D 的成本與解題率 Pareto 研究,是 AgentOpt 式的組合最佳化
  • The Verifiability Thesis — 這套設計體現「能驗證的部分就自動化」

開放問題#

  • 這套框架的適用範圍受限於 Lean 的 mathlib 成熟度。若領域需要的是新理論,而非子目標分解,該如何拓展?
  • AlphaProof 單獨使用時助益不大,作為工具則有幫助。隨著證明 LLM 能力增強,AlphaProof 工具是否會完全變得多餘?

資料來源#

§ end
Cited by 17
  • OEIS Open Benchmark×4

    Alphaproof Nexus — the system that built the item set and reported 44/492; the one whose result is…

  • Agentic Loops Overtake Bespoke Systems×3

    AlphaProof Nexus (evolution + Elo raters + AlphaProof tool) · 44 (9%) · ~$10 avg, up to ~$50 for…

  • AI-Driven Formal Proof Search×3

    The paradigm — demonstrated at research scale by Google DeepMind's Alphaproof Nexus (arXiv…

  • Evolutionary Proof Search×3

    The mechanism inside DeepMind's full-featured Alphaproof Nexus agent (agent D), inspired by…

  • FrontierMath Erdős Benchmark×3

    The motivation is that Erdős problems had become "something of a central benchmark for tracking AI…

  • Lean×3

    A proof assistant (interactive theorem prover) in which "definitions, theorems, and proofs are all…

  • Google DeepMind×2

    Google's AI research lab. In this corpus it appears as the lab behind Ai Driven Formal Proof Search…

  • Terence Tao×2

    Ai Driven Formal Proof Search, Alphaproof Nexus). The wiki is where a lab's claim becomes part of

  • Agent Harness Engineering

    Ai Driven Formal Proof Search — the Alphaproof Nexus proof-sketch-with-EVOLVE-BLOCK-markers is a…

  • Agent Loop Pattern

    Alphaproof Nexus — the framework whose basic agent (A) is a Ralph-loop fleet; it matched the…

  • Automated Conjecturing

    Alphaproof Nexus — the proof-search framework that closed a 1996 Graffiti conjecture, the result

  • Autonomous Scientific Discovery

    The survey's §5.1 calls mathematics and computer science "some of the clearest evidence of the…

  • Epoch AI

    44/492 (9%) for DeepMind's AlphaProof Nexus at comparable cost per solve.

  • Kernel-Level Proof Auditing

    for Alphaproof Nexus and Lean 4.29.1 / Mathlib 5e932f97 for ProofEvolve), and whether those

  • Many-Agent Proof Harnesses

    Every other proof-search system in this wiki — Alphaproof Nexus, ProofEvolve, LeanMarathon, the…

  • Entities — People, Orgs, Tools & Projects

    Alphaproof Nexus — DeepMind framework for LLM-aided Lean proof generation; four agents…

  • Open Questions Backlog

    Alphaproof Nexus ×2 (oldest 129d) — The framework's reach is gated by Lean's mathlib maturity.…

Related articles
  • AI-Driven Formal Proof Search

    LLM writes Lean, the compiler checks every step → no hallucination; DeepMind: 9/353 Erdős + 44/492 OEIS open problems;…

  • Lean

    Proof assistant whose compiler mechanically verifies every step; the `sorry` placeholder enables proof sketches; mathli…

  • Agentic Loops Overtake Bespoke Systems

    DeepMind's *basic* Ralph-loop agent matched its bespoke evolutionary+AlphaProof system as the LLM improved; the bitter…

  • Many-Agent Proof Harnesses

    The unformalized branch of machine proof: many-agent pipelines that write research-level proofs in natural language and…

  • Logical vs Intelligible Proof

    De Toffoli and Duede's (2026-09, `practitioner-opinion`) distinction between the *logical* notion of proof — deductive…