Sources#
- After Math
- Did OpenAI solve the wrong Navier-Stokes problem?
- FLT: Anthropic has beaten me to it
- On the Navier–Stokes Millennium Prize Problem
Summary#
On 2026-09-08 OpenAI published On the Navier–Stokes Millennium Prize Problem (On the Navier–Stokes Millennium Prize Problem, no individual byline, ~1,900 words, vendor-claim), announcing that a system of its agents, powered by an unreleased internal model, had produced a proof that three-dimensional incompressible Navier–Stokes dynamics can develop a finite-time singularity, together with a Lean formalization. A companion result on the Euler equations is announced in the same post.
This page exists because the event was already load-bearing across a dozen wiki pages before the first-party account was compiled — the "10,000 agents, 130 billion tokens, 88 hours" triple entered the corpus through Brown's Dwarkesh interview nine days later and is cited on Multi-Agent Collective Intelligence, Compute-Controlled Benchmarking, Autonomous Scientific Discovery, Evaluation Horizon Versus Release Cadence and Latent Capability Overhang. The announcement is the primary document for those figures, it adds several the interview does not, and it contradicts one thing the corpus had recorded (that no formalization was claimed). Everything below is OpenAI's claim about its own unreleased system. The tier is vendor-claim and does not move: the result is disputed, and as of 2026-09-21 no independent verification of any kind is in this corpus.
What is actually claimed#
The mathematical statement. OpenAI claims its system produced "an analytical proof and a Lean formalization that an initially smooth fluid at rest can develop a singularity in a finite time," with a smooth force applied and energy finite throughout, from rest to blow-up. In the Clay Mathematics Institute's formulation this resolves the problem by establishing statement "C" (and also "D") — the breakdown branch, not existence-and-smoothness. The mechanism described is a vortex spiralling inward and elongating ("like spaghetti"), whose central region shrinks and accelerates while total energy stays finite; the technical difficulty named is that acceleration, pressure gradient, momentum transfer and viscosity must "both become big yet cancel in a precise way," leaving the external force smooth as the velocity diverges. The figure (viewed under the two-pass rule) is a strand rendering of the same object: interleaved trajectories around a vertical axis annotated inward spiral and axial stretching, with strand colour encoding angular rotation speed — orange fast near the axis, teal slower outside. It is an illustration of the ansatz, not evidence for it.
The Euler companion, and the concession attached to it. Among "easier" problems the system was also given the regularity problem for the Euler equations (Navier–Stokes with viscosity removed), and OpenAI reports its agents resolved the unforced variant — "nearly 100 agents… approximately 50 hours" for the Euler disproof. In the "Concurrent work" section OpenAI states that Levent Alpöge (Anthropic) and Tristan Buckmaster (NYU), using an internal Anthropic model, produced a resolution of the forced Euler problem, and recognizes their priority on it: "we recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement." The forced/unforced distinction is the whole of OpenAI's claim to a distinct Euler result, and it is also its argument that "our proofs also differ significantly."
Not a prize claim. "We do not intend to claim the Millennium Prize for this result."
The run, as reported#
The timeline is unusually specific for a vendor post, and the specificity is itself the interesting part.
| Event | Date / figure (OpenAI's) |
|---|---|
| Internal model training begins | August 28 ("training is ongoing and its performance continues to improve") |
| Rumours of two resolved Millennium problems heard | Tuesday, September 1 — the trigger for the whole effort |
| Agents reach the Navier–Stokes resolution | Saturday, September 5, ~88 hours after launch |
| Lean formalization and verification | +17 hours, "via GPT‑6 Astra" |
| Full project and Lean verification complete | September 6 |
| Publication | September 8; "Concurrent work" updated September 10 |
| Euler (unforced) result | ~100 agents, ~50 hours |
| Navier–Stokes group size | "on the order of 10,000 concurrent agents" |
| Traffic, Navier–Stokes only | 2.7 million inter-agent messages, ~130 billion output tokens |
| Traffic, all attempted problems | 4.9 million messages, ~300 billion output tokens |
The model. "An internal model that is significantly more capable than GPT‑6 Astra," still in training during the run — "when a further trained version of our internal model became available over the course of the effort, we updated our agents to that model." The September 10 update adds that it "was developed through large-scale reinforcement learning on top of a previously pretrained model."
The harness, such as it is described. Agents had a cached copy of the internet and code execution; they were "subdivided into groups with the ability to communicate within the group," groups varied in size, and different groups were prompted with different variants of the problem statement — "A" and "B" (which would yield a proof) and "C" and "D" (a disproof) to separate groups, i.e. the search was run in both directions simultaneously. After the Euler result landed, resources were shifted off the other Millennium problems onto Navier–Stokes and those agents were prompted with the Euler resolution. Groups were then cross-pollinated using Codex "to consolidate the most useful insights from each agent group," drawing on the agents' own intermediate results — and OpenAI states that the group that found the Navier–Stokes solution was guided in this way. Frontier-evaluation safeguards — "monitoring and isolation" — are stated to have been in force throughout.
What this adds to the second-hand account, and the one thing it corrects#
It confirms the circulating numbers and splits them. Brown's "130 billion tokens" is the Navier–Stokes figure, not the total; the post's total across all attempted problems is ~300 billion output tokens and 4.9 million messages. The 88 hours is time-to-resolution and excludes the 17 formalization hours. So the triple in circulation is accurate as far as it goes and understates the campaign by roughly a factor of two in tokens.
It corrects the corpus on formalization. Autonomous Scientific Discovery recorded, correctly at the time, that "nothing in the account mentions formalization, Lean, or machine-checked proof; the verification is mathematicians reading the output" — because Brown did not mention it. The first-party post claims both: a Lean formalization and a verification pass, attributed to GPT‑6 Astra rather than to the internal model, at 17 hours. That claim is superseding, not corroborating, and it is the single most consequential line in the document for this wiki.
And it is where the claim is thinnest. The post says that the formalization happened and what produced it; it does not say what was formalized (the full blow-up theorem, or a skeleton with the analytic core in place?), which statement the Lean theorem states and who checked that the statement is the Millennium one, what Lean/mathlib version or axiom discipline was used, whether a sorry-free, native_decide-free, #print axioms-clean standard of the kind AI-Driven Formal Proof Search documents was applied, and what human involvement any stage had. The repo (github.com/openai/NavierStokesAndEuler) and the two PDFs (cdn.openai.com/pdf/…/navier-stokes.pdf, …/euler.pdf) are linked from the post and are not in this corpus — none was fetched at ingest. Until one is, "formalized in Lean" here is a vendor's sentence, not a kernel's verdict; the contrast with FrontierMath Erdős Benchmark, where a solve is a checker's output under a published protocol, is the cleanest one available.
Seventeen hours is also a number worth holding against the corpus's only measured formalization tax. FrontierMath Erdős Benchmark prices Erdős problem 90 — an 18-page natural-language proof from an OpenAI model — at 1.2 million lines of Lean produced by a separate, human-led effort, with the cause named as a "deep" result missing from mathlib. Navier–Stokes blow-up is not a domain where mathlib is thick either. So either an AI paid a comparable tax in 17 hours (which would be the largest single capability datum in this corpus and is asserted in one clause), or the two artifacts are not the same kind of object. Nothing in the post settles which, and the wiki should not assume the flattering branch.
The priority dispute and the data question#
The "Concurrent work" section exists because the effort was triggered by a rumour about someone else's result, which OpenAI says plainly: the September 1 rumour "we later realized was related to" Alpöge and Buckmaster. OpenAI's account of its conduct: after completing its own project and Lean verification on September 6, and believing the other team also had Navier–Stokes, it reached out to offer a concurrent release and a joint announcement recognizing their priority, learned at that point that their result was forced Euler, and offered them "visibility into all of the prompts we used and later to see the proof."
Two denials carry the weight, and they are of different strengths:
- Access. "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem."
- Training influence (this is the September 10 update, per the post's footnote 2). "Following an investigation, we have confirmed that Buckmaster's Codex prompts over the two months preceding this announcement… could not have influenced the system in any way, including through training," the stated ground being that the internal model was made by large-scale RL on an already-pretrained base.
Outside the corpus (press reporting, not a compiled source — attributed and dated here so a later compile can replace it with one): coverage on 2026-09-08 and 2026-09-10 in Quanta, Fortune, Scientific American and ABC recorded that Buckmaster alleged OpenAI adopted a strikingly similar niche method after his team's progress of August 15–22; that OpenAI said it began training the model on August 28 and "cannot rule out that de-identified data derived from their usage of our products helped improve our models"; that Charles Fefferman — author of the Clay problem description — said he was "thrilled" and credited Córdoba and Martínez-Zoroa; that Sébastien Bubeck put the compute at "several million dollars"; and that at that point no preprint, no independent Lean re-check and no refereed review existed. Two observations follow that do bear on how to read the post. First, the vendor's position hardened between the two statements — "cannot rule out" on the 8th, "could not have influenced… in any way" after the investigation on the 10th — and the investigation was OpenAI's own. Second, the August 28 training start is six days after the rival team's reported progress window, which is why the "including through training" clause is the one doing the work.
This is the corpus's first observed instance of voluntary inter-lab contact over a capability result, and it is the shape Cross-Lab Pre-Release Review describes with every variable moved: contact was post-hoc rather than pre-release, initiated by the claimant rather than the reviewer, and its object was credit rather than danger. It also reproduces that page's structural finding — there is no adjudicator anywhere. The investigating party, the accused party and the publishing party are the same.
What it is and is not evidence for#
Read as an instance of the paradigms this wiki tracks, the post supports less than its headline:
- As a many-agent result it is the largest published scale in the corpus — 10,000 concurrent agents, 2.7M messages — and it reports no baseline of any kind: no single-agent arm, no smaller-group arm on the same problem, no ablation. The nearest adjacent point is another problem: ~100 agents / ~50 hours on unforced Euler. Two problems at two scales are not a curve, and the post does not claim they are. See Multi-Agent Collective Intelligence and Many-Agent Proof Harnesses for the matched pair this forms with Google's Stellar Colosseum — one
empiricalwith a paper and a benchmark, onevendor-claimwith a Millennium problem, neither with an agent-count ablation. - As a budget disclosure it is unusually complete on tokens, messages, agents and wall-clock, and silent on dollars — the opposite balance from FrontierMath Erdős Benchmark. Compute-Controlled Benchmarking's standing point holds exactly: disclosure without a counterfactual is still uninterpretable.
- As a harness result it is a bespoke apparatus — thousands of agents, hand-designed problem-variant assignment, a mid-run model swap, Codex-mediated cross-pollination, a Lean verifier at the end — with no loop baseline. It therefore cannot settle anything in Agentic Loops Overtake Bespoke Systems, and the post itself makes no efficiency claim.
- As a capability datum about the internal model it is the second-strongest thing in the corpus after the withheld-model statements on Evaluation Horizon Versus Release Cadence, and it is the first artifact (a public writeup, PDFs and a repo) produced by a model nobody outside OpenAI can run. That is a new and weak instrument for measuring the internal/external gap: the output is public, the producer is not.
- On research taste the post says nothing that contradicts the boundary Transformative Creativity records — the problems were posed by the Clay Institute, the variants were supplied by humans, the pivot to Navier–Stokes was a human decision, and the winning group was steered.
The closing "Progress and responsibility" section is worth recording as positioning rather than content: the result is framed as "not a culmination, but rather a snapshot in time," OpenAI says it is "focusing on understanding this model," and it raises "more deliberate choices about the pace of progress" — the same pacing argument Evaluation Horizon Versus Release Cadence carries, here in the vendor's own voice on the day of its largest capability announcement.
The first compiled response: "an answer, not a solution" (2026-09-12)#
Four days after the announcement, Silvia De Toffoli (IUSS Pavia) and Eamon Duede (Princeton/Purdue) published "After Math" as a guest post on Terence Tao's blog (After Math, ~2,000 words, practitioner-opinion). It is the first response to this claim that exists in the corpus as a compiled source rather than as press reporting, and it is worth separating what it does and does not do.
What it concedes. It does not dispute the result and does not accuse anyone: "We do not deny that, if OpenAI's announcement is correct, this is an extraordinary achievement." It credits the formalization specifically — "a Lean formalization of the Navier–Stokes result is … a genuine and important contribution: by meeting the demands of the logical notion of proof, it secures certainty" — and it notes that OpenAI published two artifacts, "a Lean formalization certifying validity … accompanied by a manuscript that appears to contain the corresponding informal proof."
What it withholds, and on what ground. Its claim is that a certificate is not yet a solution, because a genuine proof must be both logically valid and intelligible — communicable in ideas a mathematician can connect to existing knowledge and build on (full treatment on Logical vs Intelligible Proof). Their verdict is explicitly provisional and dated: "At this moment, it is an answer, not a solution… Perhaps, we will find that they have [delivered a fruitful solution], but at the moment, the situation is far from clear." Their supporting witness is the Clay Institute's own page for this problem: proof matters "Because a proof gives not only certitude, but also understanding."
Why it matters for the open questions below. It is an argument about what verification would settle, so it bounds the first question's payoff rather than answering it: on this account a clean repository would establish the logical half in full and would say nothing about whether the result has entered mathematics. Note also the tier ordering — a practitioner-opinion cannot bear on whether the theorem is true, which is what the first two questions ask; it bears only on what a "yes" there would mean.
Two items of context it supplies first-hand. It quotes Tristan Buckmaster's public statement (cims.nyu.edu/~tristanb/statement.pdf) — "This is a Deep Blue–Kasparov moment." This does not replace anything in the outside-the-corpus paragraph above (which records a different Buckmaster datum, the method-appropriation allegation, and this post does not mention it); it adds a compiled one, and it complicates the press framing: the mathematician reported there as alleging appropriation is also on record calling the result epochal. (The statement PDF itself is still not in the corpus; this is De Toffoli and Duede's quotation of it.) And it reports a declaration at mathandai.org, initially signed by 25 Fields Medallists, warning of a "severe misalignment" between the goals of AI companies and those of the mathematical community — the first sign in the corpus of an organized disciplinary response to claims of this kind. Nothing else in that press paragraph is touched by this source.
The second compiled response: "the wrong problem" (Scientific American, 2026-09-21)#
Thirteen days after the announcement, Joseph Howlett reported in Scientific American (Did OpenAI solve the wrong Navier-Stokes problem?, ~1,000 words, practitioner-opinion; wording as reported, the body having been fetched through a summarizing model) a dispute that is orthogonal both to correctness and to intelligibility: whether the statement proved is the problem anyone cared about. Its argument, as reported:
- The forcing term is optional, and OpenAI's blow-up needs it. Clay's official statement (Fefferman, 2000) allows an external force in its option "C", so on that formulation the result "unambiguously solved the problem." Most researchers, per Martínez-Zoroa, "were specifically considering the scenario without a force"; Córdoba and Martínez-Zoroa had spent years on the forced case, building a "very specific external force to trigger a blowup." Luis Silvestre (Chicago): "The Clay problem is settled, but the main problem for the Navier-Stokes equations is not."
- A preprint says the method cannot close the gap. Three mathematicians posted arXiv 2609.20803 on 2026-09-17, which the article reports "showed that OpenAI's method can never be extended" to the unforced problem: remove the force and the blow-up disappears. Not in this corpus; this is a journalist's paraphrase of a paper nobody here has read.
- The open possibility. Fluid dynamicists are now considering that Navier–Stokes may blow up only with a contrived force. Cao-Labora: if so, "people would probably think about the Clay problem and say, 'We shouldn't have put the external force in the statement.'" He adds that LLMs are strong at constructing explicit examples ("finding blowups that exist") and weaker at "new theory," the reason he thinks mathematicians may be at "less of a disadvantage" on the unforced problem.
What it does to this page. Nothing above is struck: OpenAI's post states the smooth force itself ("a smooth force applied"), and this source does not dispute the proof. It reframes the headline — "resolves the Clay problem" holds, on this account, only under the reading with forcing — and it stays a statement-fidelity finding, a third axis beside Kernel-Level Proof Auditing's axiom check and Logical vs Intelligible Proof's intelligibility. It is a criticism of the Clay framing at least as much as of OpenAI, which the article says openly. The tier stays vendor-claim: this source assumes the proof valid without saying who checked it. One timing conflict is carried unresolved: the article puts OpenAI's finish "less than a day" after a rival forced-Euler blow-up on 2026-09-07, whereas OpenAI's own timeline (above) puts its Navier–Stokes resolution on 09-05 and its Lean verification on 09-06; each is a report of the other side's dates. The article also says only that the two mathematicians and the company are in "a heated dispute," and adds nothing to the training-influence question below.
Connections#
- Logical vs Intelligible Proof — the frame in which this announcement is "an answer, not a solution": Lean certifies validity without conferring the understanding the discipline wants, and the distinction is orthogonal to every dispute above about whether the proof is correct
- Terence Tao — host of that response, and the corpus's independent reference point on what machine results in mathematics amount to
- AI-Driven Formal Proof Search — the paradigm this claim would extend furthest: a vendor asserting a Millennium-Prize-level result with a Lean formalization, i.e. the kernel-checked branch reaching a problem class every measured source puts outside it (9/353, 2/68, 0/6,522). What the post does not supply is the standard — no axiom discipline, no statement review, no artifact in this corpus
- Many-Agent Proof Harnesses — the OpenAI-side counterpart of the 2026-09 matched pair of many-agent research claims: same month, same shape of claim, opposite evidence tiers, and neither with an agent-count ablation
- Multi-Agent Collective Intelligence — the corpus's largest many-agent scale datum (10,000 agents, 2.7M messages, ~130B output tokens), and its emptiest: no baseline, and the advocate's own <10% credit discount lives on that page
- Agentic Loops Overtake Bespoke Systems — a bespoke apparatus with a verifier at the end and no loop arm; the noisy-verifier and next-generation questions there both touch it and neither is settled by it
- Compute-Controlled Benchmarking — the budget is the headline and the counterfactual is absent, which is the page's own "disclosure is necessary and not sufficient" case in its most expensive form
- FrontierMath Erdős Benchmark — the disciplined opposite: a fixed $300 budget, a published item list and a kernel's verdict, run on the model this one claims to have surpassed
- Latent Capability Overhang — the withheld-capability sibling: the internal model here is claimed to be a generation past the pre-release GPT-6 Astra that scores 2/68, which widens the extracted-and-withheld gap rather than the released-and-unextracted one
- Evaluation Horizon Versus Release Cadence — the first public artifact from the internal side of the gap, and a vendor closing its announcement with the pacing argument
- Autonomous Scientific Discovery — the verification-speed spectrum this sits on, and the page this source corrects on formalization
- Cross-Lab Pre-Release Review — post-hoc inter-lab contact over credit rather than pre-release contact over danger, with the same missing adjudicator
- Transformative Creativity — the problems were posed externally and the search was steered, so nothing here crosses the problem-posing boundary
- Lean — the verifier named in the claim and absent from the evidence
- OpenAI — the claimant
- Noam Brown — the second-hand route by which these figures entered the corpus, nine days later
- Anthropic — employer of Levent Alpöge, whose forced-Euler priority OpenAI concedes
- Large-Scale Test-Time Compute (hub) — the budget axis the run sits at the extreme end of
- Verification as the New Bottleneck (hub) — the claim's whole weight rests on a verification step nobody outside OpenAI has run
- Kernel-Level Proof Auditing — the concrete audit the open question about the Lean artifact is asking for:
#print axiomson every compiled proof against the three-axiom whitelist, the only check that separates "formalized in Lean" from a proof that compiles onsorryAx; nothing in the announcement says it was run - Statement Drift — the spec-level route on that page's list: a statement faithful to Clay's forced option "C" but reportedly not to the unforced problem the field meant, a gap no comparator or implication check can see
Open Questions#
-
Does the linked Lean artifact (
github.com/openai/NavierStokesAndEuler) actually contain asorry-free, axiom-clean proof of a statement that a competent third party agrees is the Clay statement "C"? The post asserts a formalization and publishes none of the discipline; the repo and the two PDFs exist and are not in this corpus. Falsifiable by fetching and inspecting them — the cheapest open question on this page and the one everything else rests on. Partially answered 2026-09-29 (Did OpenAI solve the wrong Navier-Stokes problem?,practitioner-opinion): on the statement half, Scientific American reports that the result does resolve Clay's original formulation via its forcing option "C", which is what the post claims; but it also reports that the community reads the forced statement as a loophole rather than the intended problem, so "a competent third party agrees it is the Clay statement" now depends on which reading of Clay is meant. Thesorry/axiom half is untouched. Scope sharpened 2026-09-21 by After Math (practitioner-opinion): a clean repository would settle this question in full and would still leave the result an answer rather than a solution on De Toffoli and Duede's account — the kernel certifies deductive validity by a procedure that "does not itself require understanding," so nothing in the repository can show that the proof conveys why the theorem is true. Worth holding when reading a future "verified" headline: what the check buys is certainty, which is less than the announcement's framing implies and more than nothing. Contrast 2026-09-29 (FLT: Anthropic has beaten me to it,case-study): the Anthropic FLT repo, also a vendor artifact, was compiled and run throughcomparatorby an outside expert (Buzzard) and passed; no equivalent third-party check of the Navier-Stokes repo is reported here. -
Does the result survive independent verification? (Trigger events, any of: a refereed publication or arXiv preprint of the proof; an independent Lean re-check by a party outside OpenAI; a statement from the Clay Mathematics Institute; or a published refutation.) As of 2026-09-21 none has occurred, which is why the tier is
vendor-claim. -
How does the priority and influence dispute resolve — specifically, does any party outside OpenAI ever get to check the claim that Buckmaster's Codex prompts "could not have influenced the system in any way, including through training"? The denial rests on an investigation by the accused party, and hardened from "cannot rule out" to "could not have" in two days. (Trigger: an external audit, a disclosure of the training-data provenance, or a statement from Alpöge or Buckmaster accepting or rejecting the account.)
-
Is the unforced 3D Navier–Stokes problem in fact resolved differently, i.e. does blow-up occur without an external force, or does the forced-only scenario the article raises turn out to be the truth? (Trigger: a blow-up construction for the unforced equations, or a proof of global regularity there; also whether Clay revises the statement.)
-
What exactly does arXiv 2609.20803 prove, and does its "cannot extend" result cover only OpenAI's ansatz or every method of the forced-blow-up programme? Falsifiable by fetching and reading the paper, which this corpus reports only through Scientific American.
Sources#
- On the Navier–Stokes Millennium Prize Problem — OpenAI, "On the Navier–Stokes Millennium Prize Problem", openai.com, published 2026-09-08 with a 2026-09-10 update to "Concurrent work" (recorded in the page's own footnote 2), no individual byline, ~1,900 words,
vendor-claim— assigned and not to be upgraded: a first-party announcement about an unreleased internal model, whose result is disputed and which no third party had verified as of 2026-09-21. A web article, not PDF-derived, so nodocling:table rules apply; the body was staged from a fullbody.innerTextextract via headless browser after WebFetch andcurlboth hit a Cloudflare interstitial, and there is no selector-based under-read. The one load-bearing figure (the vortex diagram) was viewed under the two-pass rule and matches the raw's inline transcription, including the caption; three "Keep reading" thumbnails are decorative and were skipped. What is linked and not held. The proof PDF (cdn.openai.com/pdf/…/navier-stokes.pdf), the Euler PDF (…/euler.pdf) and the Lean repository (github.com/openai/NavierStokesAndEuler) are cited by the post and were not fetched — they are queued for a separate pass. Every statement about the formalization on this page is therefore a report of OpenAI's sentence about its own artifact, not an inspection of the artifact. COI is total and structural. OpenAI is the claimant, the producer of the model, the party accused of appropriating a rival's method, the investigator that cleared itself, and the publisher of the account — in one document with no named author. Attribute every number in-text. One outside-the-corpus paragraph (press reporting from 2026-09-08/10: Quanta, Fortune, Scientific American, ABC) is carried in the body under an explicit label, dated and attributed, and is not treated as vault evidence. Replace it with a compiled source when one is ingested. - After Math — Silvia De Toffoli (IUSS Pavia) and Eamon Duede (Princeton/Purdue), "After Math", guest post on Terence Tao's blog, 2026-09-12, ~2,000 words,
practitioner-opinion. The first response to this announcement in the corpus that is a compiled source rather than press reporting. No measurement, no data, no access to the artifacts — a philosophical argument about what the announcement amounts to, cited here for the answer/solution verdict, the two concessions ("extraordinary achievement" conditional on correctness; the formalization "secures certainty"), the Clay quotation, the Buckmaster "Deep Blue–Kasparov" quote and themathandai.orgdeclaration. It bears on what a verification would mean, never on whether the theorem is true — apractitioner-opinioncannot move a disputed mathematical claim in either direction. Neither Buckmaster's statement PDF nor the declaration was fetched; both are the authors' quotation. No COI: neither author has a lab affiliation or a competing result. Full treatment on Logical vs Intelligible Proof - FLT: Anthropic has beaten me to it — Buzzard, 2026-09-04,
case-study. Cited only as the contrast for artifact verification. - Did OpenAI solve the wrong Navier-Stokes problem? — Joseph Howlett, Scientific American, 2026-09-21, ~1,000 words,
practitioner-opinion. Journalism relaying named mathematicians (Silvestre, Córdoba, Martínez-Zoroa, Cao-Labora); no measurement. Fetched via WebFetch's summarizing model, so quotations are carried as reported and are not byte-verified. The central technical claim rests on arXiv 2609.20803, not fetched. Assumes the proof valid without saying who checked it; the interviewees include the forced-blow-up programme's own authors. Cited for the forcing-term loophole and its open consequences only.
Cited by 20
- AI-Driven Formal Proof Search×4
Navier Stokes Ai Claim — the claim that would put this paradigm at the top of the ladder in one…
- Transformative Creativity×4
Navier Stokes Ai Claim — the result Brown's observation is held against, now compiled first-party:…
- Agentic Loops Overtake Bespoke Systems×3
Navier Stokes Ai Claim — the apparatus axis at its extreme and the comparison axis at zero: ~10,000…
- Autonomous Scientific Discovery×3
Where the correction leaves it, and it does not move as far as it looks. A claimed kernel check is…
- Compute-Controlled Benchmarking×3
Navier Stokes Ai Claim — the first-party budget disclosure behind that headline, in four units and…
- Cross-Lab Pre-Release Review×3
openai navier stokes millennium prize solution — OpenAI (no byline), openai.com, 2026-09-08 with a…
- Latent Capability Overhang×3
openai navier stokes millennium prize solution — OpenAI (no byline), openai.com, 2026-09-08 with a…
- Lean×3
openai navier stokes millennium prize solution — OpenAI (no byline), openai.com, 2026-09-08 with a…
- Logical vs Intelligible Proof×3
Navier Stokes Ai Claim — the occasion for the post and its running example; the argument is that
- Many-Agent Proof Harnesses×3
Navier Stokes Ai Claim — the same month's other many-agent research claim, from the other lab and…
- Multi-Agent Collective Intelligence×3
Navier Stokes Ai Claim — the first-party document behind the 10,000-agent figure this page cites:…
- Noam Brown×3
openai navier stokes millennium prize solution — OpenAI (no byline), openai.com, 2026-09-08 with a…
- Open Questions Backlog×3
Navier Stokes Ai Claim: How does the priority and influence dispute resolve — specifically, does…
- OpenAI×3
Navier Stokes Ai Claim — its largest capability announcement and its most disputed: the first-party…
- Evaluation Horizon Versus Release Cadence×2
Navier Stokes Ai Claim — the first artifact published from the internal side of the gap, and the…
- FrontierMath Erdős Benchmark×2
Navier Stokes Ai Claim — the same month's undisciplined counterpart: no denominator, no budget cap,…
- Statement Drift×2
Navier Stokes Ai Claim — the spec-level case: faithful to Clay's forced option "C," reportedly not…
- Kernel-Level Proof Auditing
Navier Stokes Ai Claim — the highest-stakes unaudited instance: a vendor-claim Lean formalization…
- Formal Mathematics & Proof Search
Navier Stokes Ai Claim — OpenAI's first-party announcement (2026-09-08, vendor-claim, disputed)…
- Terence Tao
Navier Stokes Ai Claim — the claim his blog hosted the response to
Related articles
- AI-Driven Formal Proof Search
LLM writes Lean, the compiler checks every step → no hallucination; DeepMind: 9/353 Erdős + 44/492 OEIS open problems;…
- Logical vs Intelligible Proof
De Toffoli and Duede's (2026-09, `practitioner-opinion`) distinction between the *logical* notion of proof — deductive…
- Many-Agent Proof Harnesses
The unformalized branch of machine proof: many-agent pipelines that write research-level proofs in natural language and…
- OEIS Open Benchmark
Epoch AI's 492-conjecture benchmark of *open* OEIS conjectures formalized in Lean, where a model must prove or disprove…
- FrontierMath Erdős Benchmark
Epoch AI's benchmark of 68 significant *unsolved* Erdős problems — curated by Thomas Bloom from the ~652 open on erdosp…
