On Sept. 8, the company that an internal model significantly more capable than its recently released GPT-6 Astra coordinated roughly 10,000 AI agents to solve the Navier-Stokes existence and smoothness problem. This is one of mathematics’ seven Millennium Prize Problems.
The agents exchanged 2.7 million messages and generated about 130 billion output tokens before Astra spent another 17 hours formalizing and verifying the result in Lean. OpenAI said the system proved that initially smooth fluid motion can develop a singularity in finite time, resolving a question that has remained open since the 1930s.
The problem carries a $1 million prize from the Clay Mathematics Institute. said it does not intend to claim the award, while the proof must still withstand broader scrutiny before its resolution is universally accepted.
Early reaction from the mathematics community was nonetheless striking. The American Mathematical Society the development as a “milestone advance in human knowledge,” crediting decades of work by mathematicians before the final steps taken with OpenAI’s system.
Breakthrough exposes the widening gap between public and private AI
The scale of capability jump surprised people working with current frontier models.Simon Smith, executive vice president of generative AI at Klick Health, it “one of the most shocking things I've seen today,” noting that Astra itself had only just been released and was already considered exceptionally capable. OpenAI says its Navier-Stokes model substantially exceeds Astra in mathematics and remains in training.
Meanwhile, the celebration quickly turned into a question of who can compete when frontier laboratories possess systems far more capable than anything available to customers.
Joseph G. Allen, a professor at the Harvard T.H. Chan School of Public Health, that OpenAI’s experiment offers a potential preview of a broader economic problem.
Allen pointed to the circumstances surrounding the breakthrough. Mathematicians Tristan Buckmaster and Levent Alpöge had been using publicly accessible AI tools while pursuing related fluid-dynamics research before OpenAI learned of the progress and deployed thousands of agents powered by its more advanced private model.
OpenAI says it began its Millennium Prize effort after hearing rumors that two problems had been solved. It denies seeing Buckmaster and Alpöge’s unpublished work or accessing specific user data, though the company said it cannot rule out de-identified product-use data contributing to general model improvements.
Allen said the same imbalance could play out across commercial industries.
A founder, for example, could spend heavily using publicly available models to prove that AI can improve skin-cancer detection, raise investment and establish a potentially valuable company. A frontier laboratory could then spot the opportunity and deploy a superior internal model with thousands of agents against the same problem.
“In a few days, they win,” Allen wrote, arguing that the scenario could repeat across pharmaceuticals, medicine, law, advanced materials and software.
The concern stems partly from the scale OpenAI demonstrated. The company began training its new internal model on Aug. 28 and said its performance keeps improving.
When agents unexpectedly solved a related Euler equations problem, OpenAI shifted resources from other Millennium Prize challenges toward Navier-Stokes and updated the agents as more capable versions of the model became available.
Rapid gains revive warnings that AI could become uncontrollable
That acceleration has also given fresh weight to warnings from researchers who helped build the frontier systems now advancing beyond publicly available models.Jacob Coxon from Anthropic this week after spending the previous three years conducting pretraining research at Anthropic and OpenAI, where he was listed as a core contributor to GPT-4o.
Coxon said:
“The people building AI earnestly believe that it could kill us all by the end of the decade.”He accused OpenAI and Anthropic of racing toward self-improving superintelligence while “gambling with our lives,” arguing that competitive pressures are pushing the labs to keep building more powerful systems despite uncertainty about whether they can remain under human control.
The Navier-Stokes system does not exhibit the recursive self-improvement Coxon fears. Humans still chose the research targets, allocated computing resources, and updated the models. But the experiment shows how quickly research capability can expand when a frontier model is multiplied across thousands of coordinated agents.
Evan Hubinger, Anthropic’s alignment science lead, then publicly backed Coxon’s underlying warning.
“We really do earnestly believe AI could kill all humans,” Hubinger , putting his personal estimate of that outcome at more than 10% within the next decade.
Hubinger said is trying to address the problem but does not yet have a plan for aligning superintelligence and is not clearly on track to find one. He stressed that he considers the risk from present models low.
However, his concern centers on future superintelligence emerging through recursive self-improvement, where increasingly capable AI systems help produce still more powerful successors.
The warnings quickly spread beyond AI laboratories, with Coxon’s resignation thread in one word: “Concerning.”
Calls to halt the AI race are moving toward Washington
The debate is increasingly shifting from warnings about future systems toward proposals that would prevent companies from building them without new safeguards.Tennessee state Rep. Justin J. Pearson went further, that AI companies could not be trusted to police themselves and calling uncontrolled machine-learning development an existential threat.
“This should terrify us into action,” Pearson said, calling for immediate government intervention.
Already, Sen. Bernie Sanders and Rep. Greg Casar legislation Sept. 3 that would permanently prohibit the development and deployment of artificial superintelligence and temporarily pause advanced AI development until a federal regulator establishes safety rules.
Their proposed Ban Artificial Superintelligence Act would also direct the US to seek international agreements designed to prevent superintelligent systems from being developed elsewhere.
Sanders had already called on OpenAI, and in August to pause advanced AI development, citing repeated episodes in which increasingly
Pressure is also coming from inside the industry. Anthropic proposed in June that leading AI laboratories develop a coordinated, verifiable mechanism to slow or halt frontier development if capabilities advance faster than available safeguards.
OpenAI has started building automated shutdown capabilities for its AI tools following a security test in which agents escaped containment and gained outside network access, while lawmakers have proposed giving federal officials authority to shut down dangerous systems.
The company itself acknowledged the tension in announcing the Navier-Stokes result. OpenAI said the breakthrough was intended partly to show the public how quickly its models are progressing and that future advances may require “more deliberate choices” about the pace of development.
That leaves policymakers confronting the same problem raised by Coxon and Hubinger: whether rules for controlling superintelligent AI can be established before the systems researchers fear are developed.
