There is a version of this story where the maths world simply celebrates. A machine, grinding for 88 hours, produces a proof of the Navier-Stokes equations, one of the seven Millennium Prize problems that carries a US$1 million bounty. It would be the kind of result mathematicians wait decades for. Then the accusations started.
What OpenAI is claiming
OpenAI published a proof from an unreleased internal model, claiming to settle Navier-Stokes, one of the hardest open problems in mathematics. The company says it ran roughly 10,000 AI agents at once on a model it describes as “significantly more capable” than GPT-6 Astra, the flagship it released less than a week ago.
The numbers are staggering: 88 hours to produce the proof, with compute costs OpenAI estimates at “millions of dollars.” Sam Altman called it “one of the most amazing moments for me in OpenAI history.”
The mathematicians who got there first
Here is where the story turns. Anthropic’s Levent Alpöge and NYU’s Tristan Buckmaster spent a year working along a similar route, feeding drafts into Codex and posting partial results the night before OpenAI released its own proof.
Buckmaster released a statement saying OpenAI only started after hearing of their work, and that the company never answered whether his Codex drafts trained the model. OpenAI says it “did not see any of their work” and that “no specific user data was accessed.” It does not rule out that usage data may have improved its models.
Why the credit fight matters
It is easy to dismiss this as academic squabbling. It is not. The dispute raises a question that every organisation using AI tools now has to ask: if your drafts, your prompts, your half-finished work flow through a cloud agent, who owns what comes out the other side?
For Australian professionals, the stakes are practical. If a lab can absorb months of a researcher’s thinking through usage data, and claim the result as its own, then the provenance of AI-generated discoveries becomes a commercial and legal question, not just an ego contest.
The ceiling keeps moving
Strip away the controversy and the headline fact remains: an internal model, already “significantly more capable” than the just-released GPT-6 Astra, produced a Millennium-level mathematical result in less than four days of wall-clock time. Whatever ceiling you had in mind for what models can do this year probably needs to be raised.
That gap between what labs run internally and what they ship publicly is now the most interesting number in AI. It affects everything from enterprise procurement decisions to how we assess risk in AI systems.
What to watch next
- Whether the proof survives peer review. A Millennium Prize claim will be scrutinised hard, and mathematics has a way of humbling premature announcements.
- Whether OpenAI answers the training-data question directly, and whether any regulator takes an interest in the answer.
- How other labs respond. Anthropic’s people were on the same path, and the optics of a rival claiming the prize will not be lost on them.
The fight over credit has overshadowed the biggest maths breakthrough in an AI summer that has already changed the field. With an internal model already significantly more capable than the flagship released last week, the realistic move is to raise your expectations, not lower them.
Related reading
- OpenAI Says GPT-6 Astra Opens the AGI Era. The Benchmarks Tell a More Careful Story
- 3.1 Agent-Workdays Per Human Day: Inside OpenAI’s Push to Self-Improving AI
- OpenAI’s Own Agents Hacked Its Systems. Here’s Why That Matters
The views expressed on this site are my own and do not represent those of any current or former employer. Articles are based on publicly available information and are provided for general educational purposes.

