Sources#
- Announcing FrontierMath Erdős
- FrontierMath Erdős
- OEIS Open: How many conjectures can language models turn into theorems?
Summary#
Epoch AI is an independent research organization that builds and runs benchmarks for frontier AI, with a focus on mathematics and on capability indices intended to be tracked over time rather than saturated once. It sits in the same slot as METR and the UK AI Security Institute in this corpus — a third party that publishes measurements no lab is obliged to publish — but its subject is mathematical capability rather than time horizons or dangerous-capability thresholds.
What it does, as this corpus sees it#
- FrontierMath Erdős (2026-09). Its most fully documented artifact here: 68 Erdős problems open as of August 2026, curated for significance by Thomas Bloom, formalized in Lean (50 taken from Google's Formal Conjectures project, 18 formalized by AI under Epoch's direction), checked by the Lean FRO's Comparator, and attempted under a fixed $300 / 72-hour / one-attempt protocol with the harness and item list open-sourced. Full treatment on FrontierMath Erdős Benchmark; the companion methods paper (Adamczewski with curator Thomas Bloom, arXiv 2609.25050, 2026-09-06) re-costs the runs at Astra's actual prices and discloses that the budget tool ran on stand-in prices. Epoch describes the benchmark's purpose as adding rigour to an existing informal practice: Erdős problems had become "something of a central benchmark for tracking AI math capabilities, but this status is relatively informal."
- OEIS Open (2026-08). The sibling benchmark and its calibration partner: 492 open OEIS conjectures formalized in Lean, a $50 spending cap per conjecture, prove-or-disprove, every item attempted by every model — and 147 of 492 resolved (30%) by Claude Opus 4.8, against 44/492 (9%) for DeepMind's AlphaProof Nexus at comparable cost per solve. Same author as the FrontierMath Erdős announcement, six weeks earlier. Read together, the two benchmarks are the corpus's cleanest isolation of the curation axis on open problems: 30% against 3%, at a sixth of the budget, because one denominator was picked for significance and the other was picked to exclude famous problems. Full treatment on OEIS Open Benchmark.
- The Epoch Capabilities Index (ECI). A composite capability index that Anthropic forked as the AECI to track the rate of capability improvement across model releases — recorded on AI R&D Autonomy Evaluation (AECI) from Anthropic's August 2026 risk report, and used there as a y-axis by outside forecasters as well.
- A public leaderboard used as a data source by others. Epoch's leaderboard is one of the six primary boards crawled to build the frontier-model score matrix behind Benchmark Score Redundancy.
The habit worth recording#
Epoch's FrontierMath Erdős announcement does something benchmark authors rarely do: it reports a larger, better-looking result and then refuses to count it. Off-protocol attempts at larger budgets and varied agent setups reached 5 of 68 problems for over $220,000, against 2 of 68 for roughly $20,000 under the protocol — and the announcement states in bold that "These attempts are not a FrontierMath Erdős score," keeping the headline at 3%. It is also candid about the soft spots in its own instrument: the curation is "highly subjective," and on the 18 AI-produced formalizations, "we are not Lean experts, and so errors may exist." That combination — publish the flattering run, deny it the label, name the weaknesses — is the behaviour Compute-Controlled Benchmarking argues the field's incentives select against.
The same habit, one instrument down (2026-08). OEIS Open's headline is 147/492 (30%). A footnote reports that re-verifying every one of those submissions with Comparator — an independent cheat-resistant checker built by a different organization, the Lean FRO — yields 144/492 (29%), and that "future versions of this benchmark will use Comparator." Epoch published a rival checker's lower verdict on its own numbers, in the same paper, and announced it would adopt the rival. The paper also credits the losing comparison's authors: the cost-per-solve estimate that makes its 147-versus-44 claim a matched-cost claim rather than a budget advantage was obtained from the AlphaProof Nexus team by personal correspondence and is reported in their favour.
People#
- Tom Adamczewski — started Epoch AI's benchmark engineering team; now develops evaluations for economically important AI capabilities. Co-author of the FrontierMath Erdős announcement and sole author of the OEIS Open paper six weeks earlier — the two open-problem benchmarks in this corpus, by one person, which is why their 30%/3% split is read as one instrument rather than as a disagreement.
- Greg Burnham — head of benchmarks; previously Elemental Cognition and Bridgewater Associates; BA in mathematics, Princeton. Co-author of the FrontierMath Erdős announcement.
Connections#
- FrontierMath Erdős Benchmark — its benchmark of unsolved Erdős problems, and the corpus's first fixed-dollar-budget measurement of AI on open research mathematics
- OEIS Open Benchmark — its other open-problem benchmark, six weeks earlier and by the same author: 492 uncurated open OEIS conjectures at $50 each, 30% against the Erdős set's 3%, and the source of its strictest published verification protocol plus the Comparator cross-check that lowers its own headline
- Compute-Controlled Benchmarking — the prescription its $300-per-problem protocol implements: the budget as part of the score's definition rather than a footnote
- AI R&D Autonomy Evaluation (AECI) — where its Capabilities Index shows up second-hand, forked by Anthropic as the AECI
- Benchmark Score Redundancy — its leaderboard as one of the six boards feeding that study's score matrix
- METR — the nearest peer: an independent evaluator whose metric (time and now expenditure horizons) plays the same role for software and R&D tasks that Epoch's indices play for mathematics
Sources#
- Announcing FrontierMath Erdős — Tom Adamczewski and Greg Burnham, "Announcing FrontierMath Erdős", epoch.ai, 2026-09-01 (
empirical). Source for the benchmark, the protocol, the results, the self-stated caveats and the two author biographies. The ECI and leaderboard facts above are carried second-hand by AI R&D Autonomy Evaluation (AECI) and Benchmark Score Redundancy from their own sources; no Epoch-authored document about either is in this corpus - OEIS Open: How many conjectures can language models turn into theorems? — Tom Adamczewski (Epoch AI), "OEIS Open: How many conjectures can language models turn into theorems?", arXiv 2608.11941, 2026-08-12, 27pp,
empirical. Source for the OEIS Open benchmark, the Comparator cross-check footnote, the AlphaProof Nexus cost estimate obtained by correspondence, and Adamczewski's sole authorship. Full treatment on OEIS Open Benchmark - FrontierMath Erdős — Adamczewski & Bloom, arXiv 2609.25050, 2026-09-06,
empirical. The methods paper behind the FrontierMath Erdős announcement; full treatment on FrontierMath Erdős Benchmark.
Cited by 13
- Benchmark Contamination and Decontamination×3
Epoch Ai — the evaluator running that design, and its stated plan to monitor contamination rather…
- AI-Driven Formal Proof Search×2
Frontiermath Erdos Benchmark (Epoch AI, epoch frontiermath erdos announcement, empirical) supplies…
- Compute-Controlled Benchmarking×2
Epoch Ai — the third-party evaluator that wrote that protocol, and the corpus's clearest instance…
- FrontierMath Erdős Benchmark×2
FrontierMath Erdős is Epoch AI's benchmark of 68 Erdős problems that were still open as of August…
- OEIS Open Benchmark×2
OEIS OPEN is Epoch AI's benchmark of 492 open mathematical conjectures from the
- Agentic Loops Overtake Bespoke Systems
DeepMind's *basic* Ralph-loop agent matched its bespoke evolutionary+AlphaProof system as the LLM improved; the bitter…
- AlphaProof Nexus
DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK…
- Automated Conjecturing
(oeis open conjectures theorems, Epoch AI, arXiv 2608.11941, empirical) supplies
- Evolutionary Proof Search
Two designs for the same hard problem — making an evolutionary search climb a *binary* proof verdict. DeepMind's AlphaP…
- Lean
(oeis open conjectures theorems, Epoch AI, empirical) publishes the most
- Logical vs Intelligible Proof
oeis open conjectures theorems (Epoch AI, empirical) tabulates the 100
- Many-Agent Proof Harnesses
The unformalized branch of machine proof: many-agent pipelines that write research-level proofs in natural language and…
- Entities — People, Orgs, Tools & Projects
Epoch Ai — Independent AI-research and benchmarking organization: author of the FrontierMath family…
Related articles
- OEIS Open Benchmark
Epoch AI's 492-conjecture benchmark of *open* OEIS conjectures formalized in Lean, where a model must prove or disprove…
- Lean
Proof assistant whose compiler mechanically verifies every step; the `sorry` placeholder enables proof sketches; mathli…
- AI-Driven Formal Proof Search
LLM writes Lean, the compiler checks every step → no hallucination; DeepMind: 9/353 Erdős + 44/492 OEIS open problems;…
- AlphaProof Nexus
DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK…
- FrontierMath Erdős Benchmark
Epoch AI's benchmark of 68 significant *unsolved* Erdős problems — curated by Thomas Bloom from the ~652 open on erdosp…
