September 6 was a Sunday. At 12:45 p.m., NYU mathematician Tristan Buckmaster was asked whether he could get on a call “at any point today.” In two calls that afternoon, OpenAI researcher Sébastien Bubeck told him that an unreleased internal model had produced a roughly hundred-page proof of finite-time blowup for the forced Navier-Stokes equations. That is one of the Clay Mathematics Institute’s seven Millennium Prize Problems, with a million dollars attached. It is also the exact direction Buckmaster and his collaborator Levent Alpöge had been quietly attacking for a year. Weeks earlier they had proved forced blowup for Navier-Stokes’s closest cousin equations; the Navier-Stokes case itself was still unfinished on their desk.

By Buckmaster’s account, he said on the call that if OpenAI released its result the way it proposed, he would go public with what happened. The reply: “Why would you ruin your career?” He answered that he is an academic and asked why going public would ruin his career. The reply: “If you don’t want me to be nice, then I don’t have to be nice.” (Buckmaster’s public statement)

The three preprints and their Lean formalizations went up first; Terence Tao had already reviewed the mathematics on his blog by September 7. On September 8, Buckmaster released his four-page statement. The same day, OpenAI published “On the Navier–Stokes Millennium Prize Problem”. Bubeck called the circulating accusations “false and inflammatory” and denied asking for anyone’s removal from authorship, while OpenAI stated that “we (the researchers and the agents) did not see any of their work through any means until they released it publicly”.

It would be easy to file this as a he-said-he-said. Laid out as a timeline, it reads as something else: a working demonstration of what happens to the norms mathematics uses to assign priority when execution collapses, as it did here, from a year of human labor into days of compute.

What is actually in dispute

The Navier-Stokes equations describe how viscous fluids move; airflow simulation and weather modeling sit on top of them (NASA’s aerodynamics guide). The Millennium problem asks: starting from smooth initial data, do solutions stay smooth forever, or can they “blow up” in finite time, with velocities growing without bound so that the smooth solution cannot be continued and the equations stop functioning as a physical model?

A detail that gets missed: the official problem is broader than the folklore version. The Clay Institute’s formal problem description, written by Charles Fefferman (official PDF), lists four statements (A) through (D), and proving any one of them settles the problem. (A) and (B) are global regularity without forcing. But (C) and (D) are breakdown statements, and they explicitly permit constructing a smooth external force, the term in the equations representing outside energy input, to drive the blowup. Forced blowup is not a loophole; it is a sanctioned route to the prize. I went back to Fefferman’s text to check this. The version most people quote from memory, though, is the unforced one. My guess is that even if a forced proof holds up, mathematicians will keep arguing over whether Navier-Stokes has “really” been solved; that is a prediction on my part, not something I can source.

The forced route was opened over several years by Spanish mathematician Diego Córdoba and Luis Martínez-Zoroa. Buckmaster is blunt about this in his statement: the basic idea is entirely theirs, and “I believe Luis Martínez-Zoroa deserves a Fields Medal.” What he and Alpöge did was use large language models to push that program from rough forcing to smooth forcing. It was a strictly personal collaboration: Alpöge works at Anthropic, but neither employer was institutionally involved, and Buckmaster paid the OpenAI bills out of his own research funds. They worked with Claude, Codex (especially with GPT-5.6 Sol), and more recently Astra for about a year. On August 15 they obtained smooth-forcing blowup for the Boussinesq and 3D incompressible Euler equations, Navier-Stokes’s close cousin without viscosity. On August 22 the proof passed verification in Lean, the proof assistant that turns an argument into code a machine can check line by line. That rules out slips in the deductive steps themselves; what it cannot certify is that the formal statement matches the theorem you meant to prove.

Then they did the thing that is entirely proper under the old norms and fatal under the new ones: they sat on it. Buckmaster calls the first model-generated proof “the most horrendous I have ever read,” and the two worked around the clock rewriting it into something a human reader could follow before posting anything. That took over two weeks. The window was exactly that long.

The rumor was the payload

Because of Alpöge’s employer, the story mutated as it traveled: “Anthropic has resolved a major open problem.” By OpenAI’s own account, it heard the rumor on September 1 and started its push; by September 6 it was done. On September 3, after Alpöge received tips that information about their progress had reached OpenAI, Buckmaster wrote to a mathematician there to set the record straight. Three days later came the two phone calls.

The call details are worth putting on the record. Buckmaster says the opening framing was that the model had simply been given the problem statement, with “very little human input.” As OpenAI team members sent Bubeck corrections over internal chat during the call, the picture changed in real time: an entire team had been working on it, multiple directions had been tried, the models had first been set on easier problems including Euler, even the prompt he was shown had been written by prompting Codex, and “an insane amount of compute” had been used (TechCrunch reports roughly 300 billion output tokens). He asked when the first prompt had been sent. For some time the question went unanswered; the answer he says he finally got was that it had gone out in the past few days, after information about their work reached OpenAI.

Notice what OpenAI denies and what it does not. It denies having seen any of the pair’s actual work, and it has since acknowledged their priority on the forced Euler result and said it will not claim the Clay prize. It does not deny that the rumor is what set it in motion. Under the old academic ethics, “I never saw your manuscript” settles the matter. But priority as an institution rests on a hidden premise: knowing that a path works is worth little, because getting from there to a finished proof takes years of human labor. This week that premise failed. When execution costs a few days of compute, the one scarce input left is knowing which problem is ripe and which route goes through. Buckmaster writes that the word “forced” was “a bright red flag”: “Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement.” Once the coordinates leak, the search space collapses. A rumor carries no lemmas. It does not need to; the rumor itself is the payload.

There is an uglier structural layer. The pair’s entire draft history for the project lived in Codex sessions, on OpenAI’s own product. Buckmaster asked whether the model had accessed those sessions or been trained on them. On access, he was told the model does not look up user data. On training, he says, he got no answer. OpenAI later added that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models”. I have no evidence that any data was misused, and Buckmaster himself writes, “I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me.” But the structural fact stands on its own: your infrastructure vendor is also your competitor, and a reassurance like “the model does not look up user data” cannot be verified from outside. The rational move for a researcher is to stop putting unpublished ideas into a competitor’s cloud tools.

Tao’s non-renewable resource

Six days before this blew up, Tao posted a judgment on Mastodon (collected on his AI-views page): human-generated open problems in mathematics have, for the first time, become “something resembling a non-renewable resource.” Problems are endless; good problems, the kind you learn to recognize only through long immersion in a field, are scarce, and once one is publicly solved it permanently loses its value for training and evaluation. His worst case now reads like a script for this week: “even a rumor that someone is working on one can trigger a wave of AI effort to flatten it before that project matures,” incentives flip, researchers stop sharing promising directions, and the result would “reverse centuries of traditions of open science and do serious long-term damage to the future of the field.” As Fortune quotes him: “The indiscriminate strip-mining of open problems for solutions may destroy the ecosystem from which the next generation of mathematical techniques, problems, and practitioners would have developed.”

To be clear, Tao is not opposed to AI doing mathematics. His September 7 post engages seriously and generously with Buckmaster and Alpöge’s mathematics and judges that completing the remaining goals along this line, the unforced case among them, “looks very feasible” in the near future; he is explicit that the authors have not reached those goals yet. The two authors themselves leaned on LLMs throughout. The conflict runs between two ways of using the same tools: mathematicians using models to extend their own reach, and a lab pointing a compute fleet at a vein someone else had marked out.

Three conclusions

First, the evidentiary standard for priority has changed. The hardest thing Buckmaster holds is the August 22 Lean verification record (a date from his own statement, not an independently attested timestamp) and a public GitHub repository; OpenAI, for its part, released its full proof and Lean formalization for outside checking (OpenAI, Scientific American), though outside mathematicians have yet to work through it and the Clay Institute still lists the problem as open. Both sides can produce machine-verifiable proofs, so the only thing left to fight over is timestamps. Formal verification used to be insurance on correctness. In this dispute it picked up a second job: timestamping priority. “I got there first” will be argued from machine-checkable receipts, not from collegial trust.

Second, “polish before you post” has turned into a strategic weakness. The pair spent over two weeks out of respect for readers and for the community; the cost was handing someone else a weekend-sized window. There is no clean resolution here: race and you publish what the authors themselves call “AI slop”; take care and you risk being scooped. The community needs a form of publication that decouples “I reached this point” from “here is a readable paper.”

Third, for AI labs, this is a governance gap. OpenAI’s posture after the fact was decent: it acknowledged the pair’s priority on forced Euler and made no claim on the prize money. But by Buckmaster’s account, uncorroborated as it is, what appeared at the table first was a proposal to drop Alpöge from authorship and the question “Why would you ruin your career?” Even if every conversational detail ends up contested, an industry that wants scientists to bring their best problems to its models first has to convince scientists it will not walk off with the problem.

If this pace holds, the other open Millennium problems may genuinely fall one by one within a few years; I can’t say. What stays open is a different question: once they fall, whether anyone will still say their next good problem out loud.

References