The Navier–Stokes Millennium Prize Problem
Recorded: Sept. 9, 2026, 7:10 a.m.
| Original | Summarized |
On the Navier–Stokes Millennium Prize Problem Simon Willison’s Weblog Sponsored by: Portnox — Shadow AI is the new shadow IT. On Sept. 10, Forrester Research and Portnox share practical steps to regain AI agent visibility, access management, and policy enforcement. Register today 8th September 2026 - Link Blog On the Navier–Stokes Millennium Prize Problem (via) Impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems that have been subject to a $1,000,000 prize since May 24th, 2000. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would not be invited as a co-author due to OpenAI's competitive relationship with his employer. On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...] (We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost $15,000,000.) Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...] My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year. If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.) My new preferred hypothetical for this is: If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first? Posted 8th September 2026 at 11:55 pm Recent articles The Pelican comparison grid for Astra is pretty interesting - 4th September 2026
This is a link post by Simon Willison, posted on 8th September 2026. mathematics ai openai generative-ai llms training-data ai-ethics Monthly briefing Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments. Pay me to send you less! Sponsor & subscribe Disclosures |
An impressive result emerged from OpenAI concerning the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems contested since 2000, achieved by utilizing an unreleased model. This discovery was complicated by accusations of impropriety made by Tristan Buckmaster, an NYU mathematics professor who collaborated on related problems with Levent Alpöge, a mathematician working for Anthropic. The conflict arose because Buckmaster’s complaint accompanied a hastily published version of their results, leading to questions about the process and ownership of the work. Buckmaster and Alpöge reported that they collaborated on the problem for nearly a year, heavily utilizing systems like Claude and Codex (mainly GPT-5.6 Sol), before reaching a breakthrough. When seeking clarification regarding data access and training, they found that the model did not look up user data, although the full details of how the models were trained remain ambiguous to them. In response to their inquiry, OpenAI stated that while they could not guarantee specific user data was accessed, de-identified data derived from their usage might have contributed to improving the models, though the precise results and proofs differed significantly between the groups. OpenAI detailed their approach, noting that after hearing rumors of resolved Millennium Prize problems, they launched an evaluation effort across these and other high-impact problems using internal agents. These agents completed their resolution approximately 88 hours after launch, with additional formalization and verification taking seventeen hours using GPT-6 Astra. The process involved substantial computational resources; across all attempted problems, the agents generated four point nine million messages and used approximately three hundred billion output tokens. Specifically for the Navier–Stokes problem, the agents generated two point seven million messages and consumed about one hundred thirty billion output tokens. OpenAI further described their timeline, noting that their effort began on September 1st after circulating rumors related to Alpöge and Buckmaster. They intended to offer a concurrent release of their result but were unable to invite Alpöge as a co-author due to competitive relations with his employer. The author interprets this sequence of events as OpenAI recognizing an opportunity to demonstrate the power of their latest model following the emergence of solutions, rather than reflecting on the optics of preempting a team that had previously utilized similar tools. This situation prompts broader reflections on the intersection of AI capabilities and security, particularly concerning data usage and intellectual property in research. The author raises questions about how "improved model performance" is quantified when proprietary data is used for training. Hypothetically, the author considers the potential risk that using tools like ChatGPT to partially solve a problem might influence future models, potentially enabling another party to obtain solutions first. Furthermore, concerns are raised about the security implications of using large language models, such as questions regarding the possibility of data regurgitation or exposure of proprietary research directions to competitors through model training. |