A Modest Proposal

The announcement yesterday by OpenAI of a solution to the Millennium Prize Navier-Stokes problem has been causing a great deal of discussion among mathematicians. There’s a lot to be said about this story and its implications, most of which I’ll leave to those better informed than me (for one thing, I understand basically nothing about the problem, or its solution).

While talking to people I started to realize that one thing that really bothers me about these kinds of announcements is that they’ve been coming with no information at all about how the result was achieved. Emphasis has been on the question of whether the proof is correct or understandable, but I now think people should start asking another sort of question: exactly how was this done?

Sam Altman yesterday explained that

we tried this because there were rumors on the internet last week that Anthropic’s models had solved a millennium problem and we were curious if ours could do it too.

The rumor of this kind circulating was this one, with first appearance on September 4. Altman appears to be claiming that OpenAI solved the Navier-Stokes problem over the weekend, starting just with the information that maybe Anthropic had a solution. Altman is famous though for being “unconstrained by the truth.”

Tristan Buckmaster and Levent Alpoge had been working for a year on Navier-Stokes with some AI agent help, see here. The rumor I’ve heard is that Levent told colleagues at Anthropic that he and Buckmaster were close to a solution, and that this news somehow got to people at OpenAI, possibly including some information about the methods being used.

The new AI agents have profound implications for mathematics research and I think the math community needs to start discussing specific ways to deal with these. Here’s a modest proposal: the math community needs to adopt a version of the ethical standards of experimental science. If you are using AI agents, you can’t just give a proof (formalized or not), but need to also provide a detailed explanation of how these agents were used to get the result.

In an experimental science you cannot just publish results of an experiment, you also need to precisely describe the experimental setup that produced the results. One reason for this is reproducibility: other scientists need to be able to check the result by performing the same experiment. This is not so much of a concern in mathematics, where it is traditional to instead focus on the checkability of the proof.

There are other reasons though for having this standard in experimental science. Publishing the details of how to perform an experiment makes it possible for other scientists to adopt the same experimental method and make further progress by improving it or using it to get other results. I’d argue that it is now extremely important for the math community to see exactly what AI agent methods are being used, so that they can be adopted by others who want to try and use them.

In the Navier-Stokes case, I suggest that the community should demand that OpenAI provide documentation of exactly how they got this result. This would help others who want to use these methods, and it would also resolve questions about priority. What AI agents and training data did they use, what prompts were given to them and when was this done?

Besides issues of priority (which are going to be challenging since they involve the training data), this could also help with the problem of what’s a publishable result. If a result was achieved just by prompting an agent with “solve problem X”, it can be treated in a different way than results achieved through original human work and insight.

The new AI tools are going to change mathematics research. It’s critical that mathematicians have access to the details of what they are and how they are being used (as well as access to the tools themselves, a separate issue). One can imagine a future in which AI agents leave no role for human mathematicians (or, humans at all…). More hopefully, they will get integrated into human-driven activity to continue to make progress in exploring and understanding the world of mathematical objects (as well as the relations of such objects to those of the physical world).

This entry was posted in AI. Bookmark the permalink.

3 Responses to A Modest Proposal

  1. Michael Weiss says:

    Peter, reading between the lines of your important remarks, it seems to me that there are two general issues posed by the rudiments of this story.

    First, to what extent does an initial (yet unpublished) AI-enabled investigation by one group then become background source data to enable a second and subsequent AI-enabled investigation to leapfrog the first? Did the first group unwittingly make a “disclosure” in the global database of sources exploitable by the second group? If so, sorting out priority would be complex, as you indicated.

    Second, to what extent were either of the AI-enabled investigations critically dependent on an innovative human insight by one or a small group of mathematicians–i.e., the creative step being a creation of a new mathematical tool or approach not otherwise accessible to a LLM. In this case, surely it would be appropriate to recognize the seminal human contribution as the true seed of the discovery.

  2. Deane says:

    Peter, this is not the first time OpenAI has been unwilling to share the prompts they used. They want to be able to say that their model is so powerful that it can solve conjectures that the best mathematicians in the world haven’t been able to solve. So they refuse to admit that the model received any guidance from human beings on the proof. This is sad, because for us a human-AI collaborative proof of an outstanding conjecture is a spectacular success. As far as I can tell, OpenAI does not see it this way and wants it all for themselves.

    Here, they did show willingness to share the credit with Tristan, even though Tristan had nothing to do with their proof. However, this was clearly for ulterior reasons.

    One is that they do not want to share any credit with Anthropic at all. They want to brag not about how powerful AI is but only about how powerful OpenAI is.

    I totally agree we should demand transparency. I am not optimistic we will get it from OpenAI.

    As far as I can tell, Anthropic has stayed out of this. They allowed Levent to work this project using both internal models and a collaborator using ChatGPT. The CEO appears to understand that promoting the power of AI overall is still good marketing, even if OpenAI contributed to the accomplishment.

    When I read Tristan’s statement on Monday, my dominant reaction was sadness. It almost brought tears to my eyes. Later I by chance found myself in Tristan’s office with Tristan, Levent, and others. The two told me how they had not slept at all Sunday night. They were too intent on finishing the papers, writing the statement, and releasing them as soon as they could. Instead of being able to enjoy the satisfaction of their work, they had to focus on defending themselves from OpenAI. We are all discussing the future of mathematics, but all I saw was the toll this was taking on two human beings.

  3. George says:

    I agree with Peter regarding full transparency but not sure how this will work in practice for several reasons.

    First, the result was achieved with an internal model that is not available to the public.

    Second, OpenAI may try to evade transparency by saying it is verified in Lean, so why should we explain anything. If you provide a counterxample that disproves a conjecture, and everyone agrees and can verify that it is true, are you obliged to detail how did you get this result?

    I agree this is all uncharted teritory and academia really needs to step in and provide some guidance/standards expected when dealing with AI.

Leave a Reply

Informed comments relevant to the posting are very welcome and strongly encouraged. Comments that just add noise and/or hostility are not. Off-topic comments better be interesting... In addition, remember that this is not a general physics discussion board, or a place for people to promote their favorite ideas about fundamental physics. Your email address will not be published. Required fields are marked *