H
Howardism
Plate IIEntities中文HOWARDISM

Epoch AI

Independent AI-research and benchmarking organization: author of the FrontierMath family — including FrontierMath Erdős, 68 curated *unsolved* Erdős problems formalized in Lean and attempted under a published $300/72-hour budget — and of the Epoch Capabilities Index that Anthropic forked as AECI; and of OEIS Open, 492 open OEIS conjectures in Lean at $50 each where the same author measures 30% against FrontierMath Erdős's 3%; a third-party evaluator whose distinctive habit is publishing the off-protocol runs that would have made a better headline, refusing them the label of a score, and footnoting the cross-check that lowers its own number

Article metadata
Publication details
Published:September 21, 2026
Filed:Entity
Domain:Entities
Reading:7 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Epoch AI

Sources#

Summary#

Epoch AI is an independent research organization that builds and runs benchmarks for frontier AI, with a focus on mathematics and on capability indices intended to be tracked over time rather than saturated once. It sits in the same slot as METR and the UK AI Security Institute in this corpus — a third party that publishes measurements no lab is obliged to publish — but its subject is mathematical capability rather than time horizons or dangerous-capability thresholds.

What it does, as this corpus sees it#

  • FrontierMath Erdős (2026-09). Its most fully documented artifact here: 68 Erdős problems open as of August 2026, curated for significance by Thomas Bloom, formalized in Lean (50 taken from Google's Formal Conjectures project, 18 formalized by AI under Epoch's direction), checked by the Lean FRO's Comparator, and attempted under a fixed $300 / 72-hour / one-attempt protocol with the harness and item list open-sourced. Full treatment on FrontierMath Erdős Benchmark; the companion methods paper (Adamczewski with curator Thomas Bloom, arXiv 2609.25050, 2026-09-06) re-costs the runs at Astra's actual prices and discloses that the budget tool ran on stand-in prices. Epoch describes the benchmark's purpose as adding rigour to an existing informal practice: Erdős problems had become "something of a central benchmark for tracking AI math capabilities, but this status is relatively informal."
  • OEIS Open (2026-08). The sibling benchmark and its calibration partner: 492 open OEIS conjectures formalized in Lean, a $50 spending cap per conjecture, prove-or-disprove, every item attempted by every model — and 147 of 492 resolved (30%) by Claude Opus 4.8, against 44/492 (9%) for DeepMind's AlphaProof Nexus at comparable cost per solve. Same author as the FrontierMath Erdős announcement, six weeks earlier. Read together, the two benchmarks are the corpus's cleanest isolation of the curation axis on open problems: 30% against 3%, at a sixth of the budget, because one denominator was picked for significance and the other was picked to exclude famous problems. Full treatment on OEIS Open Benchmark.
  • The Epoch Capabilities Index (ECI). A composite capability index that Anthropic forked as the AECI to track the rate of capability improvement across model releases — recorded on AI R&D Autonomy Evaluation (AECI) from Anthropic's August 2026 risk report, and used there as a y-axis by outside forecasters as well.
  • A public leaderboard used as a data source by others. Epoch's leaderboard is one of the six primary boards crawled to build the frontier-model score matrix behind Benchmark Score Redundancy.

The habit worth recording#

Epoch's FrontierMath Erdős announcement does something benchmark authors rarely do: it reports a larger, better-looking result and then refuses to count it. Off-protocol attempts at larger budgets and varied agent setups reached 5 of 68 problems for over $220,000, against 2 of 68 for roughly $20,000 under the protocol — and the announcement states in bold that "These attempts are not a FrontierMath Erdős score," keeping the headline at 3%. It is also candid about the soft spots in its own instrument: the curation is "highly subjective," and on the 18 AI-produced formalizations, "we are not Lean experts, and so errors may exist." That combination — publish the flattering run, deny it the label, name the weaknesses — is the behaviour Compute-Controlled Benchmarking argues the field's incentives select against.

The same habit, one instrument down (2026-08). OEIS Open's headline is 147/492 (30%). A footnote reports that re-verifying every one of those submissions with Comparator — an independent cheat-resistant checker built by a different organization, the Lean FRO — yields 144/492 (29%), and that "future versions of this benchmark will use Comparator." Epoch published a rival checker's lower verdict on its own numbers, in the same paper, and announced it would adopt the rival. The paper also credits the losing comparison's authors: the cost-per-solve estimate that makes its 147-versus-44 claim a matched-cost claim rather than a budget advantage was obtained from the AlphaProof Nexus team by personal correspondence and is reported in their favour.

People#

  • Tom Adamczewski — started Epoch AI's benchmark engineering team; now develops evaluations for economically important AI capabilities. Co-author of the FrontierMath Erdős announcement and sole author of the OEIS Open paper six weeks earlier — the two open-problem benchmarks in this corpus, by one person, which is why their 30%/3% split is read as one instrument rather than as a disagreement.
  • Greg Burnham — head of benchmarks; previously Elemental Cognition and Bridgewater Associates; BA in mathematics, Princeton. Co-author of the FrontierMath Erdős announcement.

Connections#

  • FrontierMath Erdős Benchmark — its benchmark of unsolved Erdős problems, and the corpus's first fixed-dollar-budget measurement of AI on open research mathematics
  • OEIS Open Benchmark — its other open-problem benchmark, six weeks earlier and by the same author: 492 uncurated open OEIS conjectures at $50 each, 30% against the Erdős set's 3%, and the source of its strictest published verification protocol plus the Comparator cross-check that lowers its own headline
  • Compute-Controlled Benchmarking — the prescription its $300-per-problem protocol implements: the budget as part of the score's definition rather than a footnote
  • AI R&D Autonomy Evaluation (AECI) — where its Capabilities Index shows up second-hand, forked by Anthropic as the AECI
  • Benchmark Score Redundancy — its leaderboard as one of the six boards feeding that study's score matrix
  • METR — the nearest peer: an independent evaluator whose metric (time and now expenditure horizons) plays the same role for software and R&D tasks that Epoch's indices play for mathematics

Sources#

§ end
Cited by 13
Related articles
  • OEIS Open Benchmark

    Epoch AI's 492-conjecture benchmark of *open* OEIS conjectures formalized in Lean, where a model must prove or disprove…

  • Lean

    Proof assistant whose compiler mechanically verifies every step; the `sorry` placeholder enables proof sketches; mathli…

  • AI-Driven Formal Proof Search

    LLM writes Lean, the compiler checks every step → no hallucination; DeepMind: 9/353 Erdős + 44/492 OEIS open problems;…

  • AlphaProof Nexus

    DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK…

  • FrontierMath Erdős Benchmark

    Epoch AI's benchmark of 68 significant *unsolved* Erdős problems — curated by Thomas Bloom from the ~652 open on erdosp…