The context

It will come as a surprise to no one that I, as most of my colleagues, am deeply concerned by the recent developments in the use of commercial LLMs to approach mathematical problems. Not because these can prove technical lemmas faster or because they can produce Lean certificates (although they do, and it is something we must reflect upon as a community), but rather because of the behaviour of the companies behind them. Despite the inhuman amounts of capital that they mobilise to appear in the public eye as “revolutionising science”, “maybe curing cancer”, “putting AI data centres in space” or whatever they tell their investors, they are for-profit companies, currently burning through investor cash as they try to figure out a business model, and in order to survive their very red accounting books they must, at the very least, acquire a positive public image.

Increasingly, this goes through solving math problems, that they see as nothing but marketing stunts.1

In case the reader does not know what I am talking about, on September 8th, 2026, OpenAI released a statement claiming to have solved the Navier–Stokes millennium problem. This should have been a moment of celebration for the AI enthusiasts, and a moment of deep reflection for the mathematical community, who now has to decide how to incorporate these tools in its line of work. Instead, it has become an incredible PR flop, a masterclass of what not to do if you are in mathematics, and it has caused a sudden communal erosion of trust in a technology the likes of which I have rarely seen elsewhere.

The problem

The day before OpenAI, NYU Professor Tristan Buckmaster published a statement on his website, sharing a solution to (from what I understand) an easier version of the Navier-Stokes problem that he reached jointly with Levent Alpöge. Buckmaster states that the two of them had been working on the problem, independently of Alpöge’s employment at Anthropic, and that for this work they had been using both Claude and Codex, the LLM platforms of Anthropic and OpenAI respectively.

The trouble comes in the second part of Buckmaster’s statement, in OpenAI’s response and in successive online discourse, from which more troubling facts have emerged.

I will not give a full account here (again, the story has been thoroughly reported at this point), but the publicly known facts are as follows:

  • OpenAI put an entire team of mathematicians at work on the problem after they heard rumors of Alpöge–Buckmaster results;
  • OpenAI ran tens of thousands of parallel AI agents on the problem for days, totaling a token consumption that, had it been billed to a client, would have cost tens of millions of USD;2
  • OpenAI’s solution of the Clay problem is very similar to the ideas that Alpöge and Buckmaster were using;
  • OpenAI denied that their agents had access to the “user sessions” of Alpöge–Buckmaster, but acknowledged that those sessions might have been used to train their models beforehand.

What the mathematical community has asked since is whether Alpöge–Buckmaster’s results were indeed collected by OpenAI, used to train an LLM and ultimately reused by their LLMs without credit – after all, OpenAI has acknowledged that this might have happened. As far as I know, a more precise response to the question has not been made public.

Setting aside the “might have happened” phrasing3, mathematicians must now ask themselves if putting their own notes in a cloud-based LLM could lead the provider to scoop them and/or to use their ideas without attribution. Their LLM provider, tho whom they gave their notes, drafts and ideas, could after all be their competitor, now able to use these notes, drafts and ideas to run a trillion-dollar machine on their problem leveraging years of uncredited work.

On this merits, I am of the opinion that:

  1. If we thought the plagiarism machine would plagiarize all written human work in existence except ours, we have not been paying enough attention;

  2. If we think that the companies who say they want “to sell intelligence” are not intentioned to replace us, whose work consists in using our own intelligence, we have not been paying enough attention;

  3. If we worry that, now that we gave your trade secrets to our billionaire competitors, they will use that information to their advantage against us, we have truly not been paying enough attention.

If I sound harsh, I apologise. My intention here is not to mock or shame colleagues. These are powerful tools, I am the first to acknowledge that: I have used LLMs myself to parse through software documentation, write bash scripts and spellcheck my writing, and these are genuinely useful tools – I mean, these things were successful enough to produce a Lean certificate for a Clay problem, clearly there is some merit to their use in mathematics, even if it must be up to the community to decide what an acceptable use should look like.

I don’t see a world where these tools are never used again in mathematics, and I don’t think it would be realistic to wish for that much.

The solution?

Moving forward however, we must address the question of who are we trusting with our work. Even if we set aside the innumerable moral concerns that giving money to AI companies entail (the use in weapons of war, the mass surveillance, the pollution, the incredibly high and rising costs…), we can no longer set aside the risk that not owning our work tools entails. Until recently we kept our ideas in paper notebooks and only discussed them with trusted friends before they were ready to be published, precisely to mitigate the risk of being scooped: we have to come up with a sensible way to integrate LLMs into our work that handles this issue properly.

I do not have an answer to every question this entails. We cannot tell mathematicians to just run their own LLM on their university-issued Macbook, nor we can expect every mathematician to rent GPUs from Cloudfare, setup vLLM on them and serve it over the internet to a GUI on their machine. Perhaps this could be done faculty-wide, though? Could a math department set aside budget for a few gaming GPUs and ask IT to help running them? Could a university do that with their High Performance Cluster? Could a national funding institution run inference with open source models for the exclusive use of their scientists? Should it do that?

Time will tell what a solution might look like, but the need for one is, I believe, no longer questionable.

  1. The first version of this draft started to exist on September 9th, 2026, and now that every mathematician with an internet connection knows the story of Alpöge–Buckmaster and the Navier–Stokes equations, it has needed a thorough revision. ↩

  2. Let this be clear: the cost of this marketing stunt was the equivalent of a dozen ERC grants. How much good research would have been produced by a dozen ERC grants focused on the same problem instead? How many talented scientists and productive members of society would this money have formed? ↩

  3. The reader must understand how utterly insane it is that “have you put this piece of data in the training set or not?” is a question that OpenAI has been unable or unwilling to answer conclusively: if you are an AI lab allegedly worth trillions of dollars, and your entire purpose is to train LLMs, how can you not know how you trained your flagship product? ↩