OpenAI Just Cleared Navier-Stokes in 88 Hours. It Spent 22 Times What Fermat Cost Anthropic.
OpenAI published the writeup on Monday, September 8, 2026: a group of roughly 10,000 agents, running on an unreleased model the post describes as significantly more capable than GPT-6 Astra, produced an analytical proof and Lean formalization showing that Navier-Stokes dynamics permit a finite-time singularity, a spinning vortex that thins into a filament until the equations break. The run took about 88 hours from first agent launch, moved roughly 2.7 million messages between agents, and burned in the neighborhood of 130 billion output tokens. Clay Mathematics Institute lists Navier-Stokes as one of seven Millennium Prize problems. OpenAI said it will not claim the prize.
Four days earlier, Anthropic published its own receipt: dozens of Claude agents produced the first complete Lean 4 formalization of Fermat's Last Theorem in eleven days on roughly six billion output tokens. Same shape of story. Very different numbers.
Two harness receipts on hard math, four days apart, from the two labs at the top of the frontier tier. That is the story worth reading. The proofs themselves are already being litigated by working mathematicians, and one of them turned into a public credit fight before the ink was dry. But the compute delta is the piece that reprices the segment, so start there.
The Delta
Line the receipts up.
| Metric | OpenAI Navier-Stokes | Anthropic Fermat | Ratio |
|---|---|---|---|
| Wall clock | 88 hours | 264 hours (11 days) | 3x faster |
| Agent count | ~10,000 | Dozens | ~300x more |
| Output tokens | ~130B | ~6B | ~22x more |
| Messages exchanged | ~2.7M | Not published | n/a |
| Model | Unreleased post-Astra | Claude, general availability | n/a |
| Orchestrator | Not disclosed | Prove2Me DAG | n/a |
| Verifier | Lean 4 | Lean 4 | Same |
Two of those rows are the numbers that decide the market read. Agents scaled roughly 300x and tokens scaled roughly 22x, so the average agent on the OpenAI run generated far fewer tokens than the average agent on the Anthropic run. That is what you would expect from fanning wide instead of deep, and it is the operational tell that the two labs are betting on different orchestrator shapes: Anthropic on a directed graph of proof obligations with a small population of agents grinding it, OpenAI on a large swarm exploring many candidate paths in parallel and letting Lean adjudicate.
Neither shape is obviously right. What matters for the segment is that both work.
What the Run Cost
OpenAI did not publish a compute bill. Build one from public pricing to see the shape.
The model was not Astra; the post says the unreleased successor is significantly more capable. Price the run against Astra's $10 input, $50 output as the floor, since the successor is not cheaper. Output alone: 130 billion tokens at $50 per million is $6.5 million. Agentic runs at this shape carry inputs 20 to 50 times higher than outputs once cached context, tool traces, and Lean feedback are counted, with cache reads dominating. Anchor the estimate at four trillion input-equivalent tokens with 90 percent cached, price the cached portion at Sol's $0.50 per million (Astra's cache-read line was not published, so use the closer OpenAI tier as the floor), and cache reads alone land near $1.8 million. Add fresh input at $10 per million on the uncached 400 billion and you pick up another $4 million. All in, the run sits in the low-to-mid eight figures, call it $10 million to $15 million of compute at list price.
That is not a rounding error, and it is not a new hyperscaler contract either. It is somewhere between a hard-tech seed round and a Series A, spent inside 88 hours on one research question. Anthropic's Fermat run, by the same accounting shape and against Fable 5 list pricing, landed at $500,000 to $1.5 million. The delta at the receipt level is roughly ten to twenty times. The delta in wall clock is three times faster. The delta in agent count is roughly 300 times.
Buyers can read that two ways. Either OpenAI paid a ten-to-twenty times premium to compress the wall clock by three, in which case the exchange rate between compute and time on hard research is now a public number. Or the swarm architecture buys something the DAG architecture cannot, in which case the receipt says the frontier just went somewhere Anthropic's Prove2Me shape cannot follow. The public evidence is consistent with both reads, and the labs will spend the next quarter arguing about which one it is.
The Credit Fight
Two hours after the announcement, NYU mathematician Tristan Buckmaster posted a statement on his website saying he had been working the same problem for nearly a year with Levent Alpöge, a mathematician employed at Anthropic in a personal capacity outside company research, and that the two of them had arrived at a working solution on August 22 using a mix of models, primarily OpenAI's Codex. Buckmaster said OpenAI researcher Sébastien Bubeck learned of that work through professional channels and, according to the statement, offered him two options: publish a partial development with OpenAI to release its full proof the following day, or write a solo paper crediting the OpenAI model but omitting Alpöge's name because Alpöge is employed at a rival lab. Buckmaster declined both proposals. OpenAI has denied using their unpublished work. Terence Tao, in a brief note picked up by several outlets, called the situation a lament without saying more.
Treat the dispute as data about the segment, not about the proof. Three things are true simultaneously. First, the compute the OpenAI run used is not the piece that produces a proof: Fermat needed Kevin Buzzard's multi-year Lean blueprint to be tractable at all, and every honest reading of the Anthropic writeup said so. Navier-Stokes evidently needed a similar upstream plan, and a version of that plan appears to have existed inside the Buckmaster-Alpöge collaboration. Second, whether OpenAI used that plan or not, the strategic value of being first to announce is high enough that the negotiation over authorship happened at all, which tells you the labs now treat these outputs as reputational assets on the same shelf as a benchmark score. Third, the professional machinery for adjudicating who did what on a proof produced by 10,000 agents does not exist yet. Journal peer review does not scale to a Lean file with tens of millions of lines. Preprint priority conventions assume a human first author. Neither is ready to answer the question OpenAI just posed.
That is the story: the receipt is real, the proof is probably real, and the profession that would normally verify it is a few years behind the calendar.
What Repeats and What Does Not
Our Fermat writeup last Friday made a specific claim: the harness is the product, and the model layer contributes general reasoning while the orchestrator contributes the ability to remember and to parallelize. The Navier-Stokes run either confirms that thesis at a larger scale, or reveals a second axis nobody had priced yet, which is the axis of raw parallel search when the plan already exists.
Read the caveats from that piece against this week. We wrote that formal proofs at this scale still require a pre-existing human blueprint, and we named Riemann, Hodge, and Navier-Stokes as three problems without one. Four days later a proof of Navier-Stokes shipped with a credit dispute attached to it that suggests a blueprint did in fact exist, just outside the announcing lab. The caveat held; it just moved. The lab with the swarm still needed the plan. It appears to have been prepared to argue about who produced it.
Two effects on the coding and research segment fall out of this. First, the harness race just became a scale race as well. Anthropic's pitch has been that a clever orchestrator on top of a general model beats a bigger cluster; OpenAI's Monday receipt is a direct answer that a very big swarm beats a clever graph on wall clock, holding the target problem constant. Second, cached input pricing matters even more than it did on Friday: the OpenAI run is dominated by cached reads at the OpenAI cache-read rate, and Anthropic's September 1 cut of Fable cache reads to $0.25 per million is the reason a hypothetical rematch would look different. Move the delta on cache reads another 50 to 75 percent and the Fermat-shape budget lands somewhere OpenAI cannot match on Astra pricing without cutting its own line.
Our Take
The interesting sentence in the OpenAI post is the one about the model. It is not Astra. It is a successor OpenAI has not shipped, has not priced, and has not put in front of a third party. Anthropic's Fermat proof ran on the same Claude every paying customer on the API can call today; OpenAI's Navier-Stokes proof ran on a model no customer has ever touched. That is a real difference and it should be part of every read on the two receipts. Anthropic is selling the workflow underneath a shipping model; OpenAI is using a science result to preview a model it has not launched. Same segment, different product motion.
Read alongside the Pachocki essay from Saturday, the Monday post is a datapoint on cadence. Pachocki called for third-party oversight of frontier model releases; two days later his employer posted a research result on an unshipped model with no third-party verification of the compute, the training run, or the proof. Both statements can be sincere at once, and both are on the record.
Practical read for anyone building on the API. If your product needs long-horizon research shape, both harness patterns are now viable at frontier scale, and the choice between them is a cost-of-cache question more than a cost-of-model question. If you can afford to fan wide, the OpenAI-style swarm shortens the wall clock; if you cannot, the DAG-style orchestrator gets there for a tenth to a twentieth of the compute on the one public data point we have. Neither buys you the upstream plan. That still comes from somewhere else, and this week is a reminder that where it comes from is a question the professional machinery has not yet caught up to.
Three Signposts
Watch three things over the next 60 days. One, whether Buckmaster or Alpöge publishes the collaboration's working notes with dates attached, which is the direct test of the priority claim rather than the tactical one. Two, whether OpenAI ships the model the Navier-Stokes run used to any external evaluator with a signed report before the model appears on the pricing page, which is the direct test of whether Pachocki's essay reflects a policy the company will adopt without a law requiring it. Three, whether a third lab (Google, xAI, Meta, or a Chinese frontier lab) posts a hard-math receipt of its own inside the same window, which is the direct test of whether the harness race just became the default sales motion for the top tier. Any two of the three fire and the shape of the segment is set for the rest of the year.
We track the underlying model economics on our models directory and the harness thesis in our recent originals. The next receipt is coming faster than this one; the compute line has already been drawn.
