資料來源#
- Announcing FrontierMath Erdős
- FrontierMath Erdős
- OEIS Open: How many conjectures can language models turn into theorems?
摘要#
Epoch AI 是一家獨立研究組織,為前沿 AI 建立並執行基準測試,著重於數學,以及旨在隨時間追蹤、而非一次就達到飽和的能力指數。在這份語料中,它與 METR 和 UK AI Security Institute 扮演相同角色——發布各實驗室沒有義務公布的測量結果的第三方;但它研究的是數學能力,而非時間範圍或危險能力門檻。
依這份語料的觀點,它做些什麼#
- FrontierMath Erdős(2026-09)。 這裡記錄最完整的成果:截至 2026 年 8 月仍未解的 68 道 Erdős 問題,由 Thomas Bloom 按重要性精選,以 Lean 形式化(50 道取自 Google 的 Formal Conjectures 專案,18 道由 AI 在 Epoch 指導下形式化),由 Lean FRO 的 Comparator 檢查,並依固定的 $300 / 72 小時 / 僅一次嘗試規範進行測試,同時公開 harness 和題目清單。完整介紹見 FrontierMath Erdős Benchmark;配套方法論論文(Adamczewski 與策展人 Thomas Bloom 合著,arXiv 2609.25050,2026-09-06)按 Astra 的實際價格重新計算測試成本,並揭露預算工具採用的是替代價格。Epoch 說明此基準的目的,是為既有的非正式實務增添嚴謹性:Erdős 問題已成為「追蹤 AI 數學能力的核心基準之一,但這個地位相對而言仍屬非正式」。
- OEIS Open(2026-08)。 與之並列的基準及校準夥伴:492 道以 Lean 形式化的開放 OEIS 猜想,每道猜想的支出上限為 50 美元,須證明或反證,每個模型都必須嘗試所有題目——Claude Opus 4.8 解決了 492 道中的 147 道(30%);以每道題目相近的求解成本相比,DeepMind 的 AlphaProof Nexus 為 492 道中的 44 道(9%)。作者與 FrontierMath Erdős 公告相同,早了六週。合讀這兩項基準,可以最清楚地看出這份語料中開放問題的策展軸線:30% 對 3%,預算只有六分之一,因為其中一個分母是按重要性挑選,另一個則刻意排除知名問題。完整介紹見 OEIS Open Benchmark。
- Epoch Capabilities Index(ECI)。 一項綜合能力指數,Anthropic 分叉採用為 AECI,用來追蹤不同模型版本之間的能力進步速度——Anthropic 於 2026 年 8 月風險報告中的 AI R&D Autonomy Evaluation (AECI) 有記錄,外部預測者也在該處將它用作縱軸。
- 作為他人資料來源的公開排行榜。 Epoch 的排行榜是建構前沿模型分數矩陣時所爬取的六個主要排行榜之一;該矩陣見於 Benchmark Score Redundancy。
值得記錄的作風#
Epoch 的 FrontierMath Erdős 公告做了一件基準作者鮮少做的事:它公布了一個規模更大、看來更漂亮的結果,接著拒絕將它計入。 在規範之外使用更高預算和不同代理程式設定進行的嘗試,以超過 220,000 美元解出 68 題中的 5 題;依規範測試則以約 20,000 美元解出 68 題中的 2 題。公告以粗體聲明:「這些嘗試不算 FrontierMath Erdős 分數」,因此標題仍維持 3%。公告也坦率指出自身測量工具的不足之處:策展「高度主觀」,而對於 18 道由 AI 產出的形式化內容,「我們並非 Lean 專家,因此可能存在錯誤。」這種作法——公布亮眼的測試結果、拒絕給它相應的標籤、指出自身弱點——正是 Compute-Controlled Benchmarking 所主張會被領域誘因排斥的行為。
同樣的作風,再往下一層看一項工具(2026-08)。 OEIS Open 的標題數字是 147/492(30%)。一則註腳報告指出,以 Comparator——由另一個組織 Lean FRO 建立、可防作弊的獨立檢查器——重新驗證所有提交後,結果是 144/492(29%),並表示「此基準的未來版本將採用 Comparator」。Epoch 在同一篇論文中公布了另一個檢查器對自身數字較低的判定,並宣布將改採對方的工具。論文也肯定了比較中落敗一方的作者:讓 147 對 44 能成為同成本比較、而非預算優勢的每題求解成本估算,是透過個人通信向 AlphaProof Nexus 團隊取得,且報告時對他們有利。
人物#
- Tom Adamczewski — 創立 Epoch AI 的基準工程團隊;目前為具經濟重要性的 AI 能力開發評估方法。他是 FrontierMath Erdős 公告的共同作者,也是六週前 OEIS Open 論文的唯一作者——這份語料中的兩項開放問題基準出自同一人之手,因此這兩者 30% 與 3% 的差距會被視為同一測量工具的表現,而非彼此意見不合。
- Greg Burnham — 基準測試主管;曾任職於 Elemental Cognition 和 Bridgewater Associates;擁有 Princeton 數學學士學位。FrontierMath Erdős 公告的共同作者。
相關連結#
- FrontierMath Erdős Benchmark — 其未解 Erdős 問題基準,也是這份語料中首次以固定美元預算測量 AI 解決開放研究數學問題的基準
- OEIS Open Benchmark — 另一項開放問題基準,早六週發表,作者相同:492 道未經策選的開放 OEIS 猜想,每題 50 美元;求解率為 30%,Erdős 題組則為 3%;也是其最嚴格的已發布驗證規範,以及降低自身標題數字的 Comparator 交叉檢查結果之來源
- Compute-Controlled Benchmarking — 其每題 300 美元規範所實踐的主張:預算是分數定義的一部分,而非註腳
- AI R&D Autonomy Evaluation (AECI) — 其 Capabilities Index 在此以第二手資料出現,並由 Anthropic 分叉採用為 AECI
- Benchmark Score Redundancy — 其排行榜是提供該研究分數矩陣的六個排行榜之一
- METR — 最相近的同儕:獨立評估者,其指標(時間以及如今的支出時間範圍)對軟體與 R&D 任務所扮演的角色,與 Epoch 的指數對數學所扮演的角色相同
資料來源#
- Announcing FrontierMath Erdős — Tom Adamczewski 和 Greg Burnham,「Announcing FrontierMath Erdős」,epoch.ai,2026-09-01(
empirical)。基準、規範、結果、自述的限制,以及兩位作者的簡介皆以此為來源。上文的 ECI 和排行榜資訊是由 AI R&D Autonomy Evaluation (AECI) 和 Benchmark Score Redundancy 根據各自來源轉述;這份語料中沒有 Epoch 撰寫的相關文件 - OEIS Open: How many conjectures can language models turn into theorems? — Tom Adamczewski(Epoch AI),「OEIS Open: How many conjectures can language models turn into theorems?」,arXiv 2608.11941,2026-08-12,27 頁,
empirical。OEIS Open 基準、Comparator 交叉檢查註腳、透過通信取得的 AlphaProof Nexus 成本估算,以及 Adamczewski 唯一作者身分皆以此為來源。完整介紹見 OEIS Open Benchmark - FrontierMath Erdős — Adamczewski 與 Bloom,arXiv 2609.25050,2026-09-06,
empirical。FrontierMath Erdős 公告背後的方法論論文;完整介紹見 FrontierMath Erdős Benchmark。
Cited by 13
- Benchmark Contamination and Decontamination×3
Epoch Ai — the evaluator running that design, and its stated plan to monitor contamination rather…
- AI-Driven Formal Proof Search×2
Frontiermath Erdos Benchmark (Epoch AI, epoch frontiermath erdos announcement, empirical) supplies…
- Compute-Controlled Benchmarking×2
Epoch Ai — the third-party evaluator that wrote that protocol, and the corpus's clearest instance…
- FrontierMath Erdős Benchmark×2
FrontierMath Erdős is Epoch AI's benchmark of 68 Erdős problems that were still open as of August…
- OEIS Open Benchmark×2
OEIS OPEN is Epoch AI's benchmark of 492 open mathematical conjectures from the
- Agentic Loops Overtake Bespoke Systems
DeepMind's *basic* Ralph-loop agent matched its bespoke evolutionary+AlphaProof system as the LLM improved; the bitter…
- AlphaProof Nexus
DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK…
- Automated Conjecturing
(oeis open conjectures theorems, Epoch AI, arXiv 2608.11941, empirical) supplies
- Evolutionary Proof Search
Two designs for the same hard problem — making an evolutionary search climb a *binary* proof verdict. DeepMind's AlphaP…
- Lean
(oeis open conjectures theorems, Epoch AI, empirical) publishes the most
- Logical vs Intelligible Proof
oeis open conjectures theorems (Epoch AI, empirical) tabulates the 100
- Many-Agent Proof Harnesses
The unformalized branch of machine proof: many-agent pipelines that write research-level proofs in natural language and…
- Entities — People, Orgs, Tools & Projects
Epoch Ai — Independent AI-research and benchmarking organization: author of the FrontierMath family…
Related articles
- OEIS Open Benchmark
Epoch AI's 492-conjecture benchmark of *open* OEIS conjectures formalized in Lean, where a model must prove or disprove…
- Lean
Proof assistant whose compiler mechanically verifies every step; the `sorry` placeholder enables proof sketches; mathli…
- AI-Driven Formal Proof Search
LLM writes Lean, the compiler checks every step → no hallucination; DeepMind: 9/353 Erdős + 44/492 OEIS open problems;…
- AlphaProof Nexus
DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK…
- FrontierMath Erdős Benchmark
Epoch AI's benchmark of 68 significant *unsolved* Erdős problems — curated by Thomas Bloom from the ~652 open on erdosp…
