On September 9, Andreas Thom, a group theorist at TU Dresden, posted his correspondence with an OpenAI researcher on Mathstodon. In early August, OpenAI had announced that an internal model solved ten open problems in mathematics and theoretical computer science, and the core of one proof happens to build on work by Thom and Gábor Kun. Thom emailed the researcher with two questions. Did his conversations with ChatGPT over the past several months, in which he discussed exactly this mathematics, enter the training data? And could the system that produced the proof access those conversations?

The researcher, Mark Sellke, answered in one line: “Regarding your conversations with ChatGPT: that did not happen.”

A month later, OpenAI published an official statement in the middle of a much larger dispute, and one sentence in it supplied, in effect, the honest answer to Thom’s first question: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

Thom is the second mathematician in a week to demand answers from OpenAI in public. Start with the first.

One week, two mathematicians

The Navier-Stokes equations describe how fluids move. The corresponding Millennium Prize problem, posed in 2000, asks whether solutions starting from smooth initial data can “blow up” in finite time, meaning the solution loses the smoothness the problem requires and the equations break down as a description of the flow. The Clay Mathematics Institute’s official problem statement offers four alternative formulations, and proving any one of them counts; the two “blowup” versions allow a smooth external forcing term in the equations.

Tristan Buckmaster of NYU and Levent Alpöge, a mathematician at Anthropic, spent close to a year on this blowup route as a personal collaboration, working with AI models including Claude and OpenAI’s Codex. The route’s basic idea comes from a forced-blowup program that Diego Córdoba and Luis Martínez-Zoroa had been developing for years. In his public four-page statement, Buckmaster writes that almost nobody else he knew of was working on it, and that the pair’s entire draft history for the project lived in Codex sessions. The statement records the timeline (see also TechCrunch’s report): on September 3 he told a mathematician at OpenAI that the pair had results they were confident in and would release soon, making clear it was a personal effort with no institutional involvement from either employer. Three days later, OpenAI researcher Sébastien Bubeck told him on a call that an internal model had produced a roughly hundred-page proof of finite-time blowup for forced Navier-Stokes, along the same route. The statement further alleges that OpenAI proposed he publish the result as sole author, dropping Alpöge, who works at Anthropic, from the byline; and that when he said he would go public, he was asked, “Why would you ruin your career?” Bubeck denied the authorship allegation and publicly acknowledged the pair’s priority.

On September 7, the two got ahead of OpenAI and released finite-time blowup proofs, with smooth forcing, for three neighboring equations: the incompressible porous medium equation, the two-dimensional Boussinesq equation, and the three-dimensional incompressible Euler equations. By the statement’s account, this took Córdoba and Martínez-Zoroa’s blowup constructions, originally built with rough forcing, and used LLMs to push them to smooth forcing and to extend them to Euler and the other equations. Terence Tao introduced the work on his blog the same day, noting that the arguments were heavily AI-assisted and that the methods look likely to extend to Navier-Stokes itself. OpenAI went on to announce a solution to forced Navier-Stokes blowup anyway, saying it would not claim the million-dollar prize. Alpöge’s response to OpenAI’s “cannot rule out” sentence was sardonic: “i mean props to them for straight coming clean.”

Thom’s storyline starts earlier. In 1999, Gromov asked whether every group, the algebraic structure that describes symmetry, can be well approximated by finite permutations. Groups that can be are now called sofic groups, and in 27 years nobody had produced a counterexample. Among the ten results OpenAI published on August 1 is the first construction of a non-sofic group, and a key step of the proof builds directly on a 2016 theorem of Kun’s and the 2019 Kun–Thom paper. According to the Northeast Times, the original writeup of the result did not mention the two mathematicians’ contribution; the credit was quietly added after mathematicians pointed it out. The result is already generating follow-up work in the community: a mathematician has since constructed a torsion-free non-sofic group in response.

What made Thom suspicious was the choice of route. As he describes it in his Mathstodon post, the Kun–Thom approach was never the leading candidate; the community had put its money on a different route through quantum computational complexity. Meanwhile, he had spent months discussing exactly this territory with ChatGPT: expander matching problems (expanders are sparse graphs with unusually strong connectivity) and extensions of the Kun–Thom work. The model walked precisely this unfashionable path to the finish. Thom switched off the “improve the model for everyone” setting on his account on June 29, but per OpenAI’s own documentation, that control only governs future conversations. Whether earlier conversations went into training, OpenAI does not say, and nothing in its public data controls (data controls FAQ) offers a per-conversation record of what entered training, so a user has no way to check.

Three denials, three scopes

Let me state my position first: there is no public evidence that OpenAI took either group’s work (Axios’s overview likewise stops at open questions). Buckmaster is blunt in his statement: “I am not accusing anyone of anything.” He recounts only what he was told, when, and what was proposed to him; Thom, too, is raising questions rather than making claims. Buckmaster’s core doubt is the coincidence of timing and route. Thom’s core doubt is that nobody would answer him directly. What deserves dissection is the structure of OpenAI’s three denials.

Sellke told Thom “that did not happen.” Thom asked two questions; on the most natural reading, that sentence covers only the second one, whether the proof-producing system accessed his conversations. The training-data question, as of publication, OpenAI has not publicly answered.

OpenAI’s official statement on the Buckmaster affair says: “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” What is being denied is that humans saw the drafts, and that the system retrieved a specific user’s data at solve time. Both are narrow enough to deny cleanly.

Only then, in the same statement, comes the sentence that matters: it cannot be ruled out that de-identified usage data improved the models. That is the layer the mathematicians were actually asking about, and at that layer the answer is that it cannot be ruled out.

Put the three sentences side by side and the pattern I read is this: each denial precisely covers the narrowest scope that can be denied, and the widest scope stays blank until public pressure forces it out. I don’t think this is a public relations blunder. My guess is that at the training-pipeline layer, OpenAI itself may not be able to give a clean answer: nothing in the public record suggests any mechanism that can trace a given conversation to a given training batch.

Tools built for personal privacy won’t protect an unpublished theorem

It is tempting to read this dispute as a privacy story, but the toolkit built for personal privacy does little to protect unpublished results, for three reasons that sit at different depths.

The first is a mismatch in what gets protected. The privacy-compliance toolkit (de-identification, anonymous aggregation, opt-outs) protects identity. Strip Thom’s name out of a conversation and the mathematical ideas survive intact; academic priority lives in the content itself and has nothing to do with identity. This is the sharpest point in Thom’s argument: protections designed for personal data do not touch the problem of ideas being learned.

The second is the direction of time. The training toggle governs future conversations only; it does not retroactively pull earlier ones out of training. And no user-visible record ties a given conversation to a given training batch, so there is no way to reconcile accounts after the fact. In the Hacker News discussion, which had over 600 comments as of publication, several users also report the toggle resetting to “on” without their knowledge. I have not verified those reports, but their direction is beside the point. What matters is that the mechanism’s entire credibility rests on the vendor’s unilateral execution, with no verifiable receipt given to the user.

The third is the most fundamental: self-investigation cannot be falsified. Academia handles the same conflict of interest with institutions. Journal referees are bound to confidentiality and barred from using a manuscript for their own advantage, misuse carries professional consequences, and a third party, the editorial board, adjudicates disputes. AI companies now occupy a structurally identical seat: tool supplier to researchers, and competitor on the very same problems. The counterpart machinery is zero. Nothing OpenAI has made public offers a traceable record of training-data provenance or model checkpoints open to third-party audit, and in this dispute every investigation and every statement so far has come from OpenAI itself. It is as if a journal referee could decide whether to scoop a submission, and then, if trouble followed, write the investigation report personally.

What researchers can do now

For people doing research with these tools, I can offer two concrete judgments.

Manage the conversations in a consumer account as “already disclosed to a potential competitor.” Enterprise and API customers get contract-level no-training commitments; I’ve previously taken apart OpenAI’s zero-data-retention policy. Individual accounts get none of that by default. One suggestion from the HN thread I agree with completely: universities and research institutes should negotiate contract terms the way enterprise customers do, instead of leaving each researcher alone with a switch they cannot audit.

Let priority rest on timestamps, not on a perfect manuscript. Buckmaster and Alpöge posted a far-from-polished version first, on September 7, and their priority was publicly acknowledged by the other side. Buckmaster apologized in his own statement for the state of those writeups. In an environment where AI companies are also racing on the problems, the cost structure of “save it until the writeup is beautiful” has flipped: a rough version posted early is worth more than a polished version posted late.

As for the AI companies’ side, Tao’s criticism operates at a different level. He said publicly that AI companies are treating these long-standing problems as marketing material for model capability, and warned that indiscriminate strip-mining of open problems for solutions may destroy the ecosystem that future mathematical techniques grow out of. Follow that criticism one step further. Gromov’s question stood for 27 years, Navier-Stokes for 26; this stock of open problems was accumulated by generations of mathematicians. Once AI tools start consuming the stock at speed, the scarce skill shifts from solving problems to posing them: judging which new questions are worth attacking and stating them in a form one can actually work on. That is something I do not yet see models doing in anyone’s place. My own judgment is narrower. What this week actually consumed is the credibility of AI companies as participants in science. Mathematics keeps careful score on priority, and in this community a vague denial gets expensive fast. Mathematicians are simply the first to tally the bill. The same structural conflict is waiting for every field that feeds unpublished ideas into closed tools.

References