AI NewsOpenAIās math solutions arenāt meeting the fieldās standards yet
OpenAIās math solutions arenāt meeting the fieldās standards yet
3:04 AM IST Ā· October 9, 2026

When OpenAIreleasedhundreds of claimed solutions to some of the worldās hardest math problems this week, the frontier lab said that it had consulted an advisory group of elite mathematicians to avoidthe controversythat came with the last time one of its models solved a long-standing problem in the field. But OpenAI fell short of those standards, particularly where the mathematicians emphasized the need for human understanding of a mathematical result. Thatās especially concerning after a new paper highlighted gaps between the natural language and formally expressed solution to a million-dollar problem ostensibly solved by OpenAIās models. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton Universityās Institute for Advanced Studies, is made up of nine prominent researchers at institutions around the world. The organization releasedguidelinesfor frontier labs solving math problems at the end of September. In a statement on the latest set of proofs, the AGMAI said that āit is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.ā However, the organizationās first request was āto stop testing advanced mathematical problems on proprietary models.ā OpenAIās release explicitly says that it is evaluating its proprietary models using open research problems in mathematics. The advisory group did not respond when asked by TechCrunch for a more thorough evaluation of OpenAIās latest proof release. The lab clearly followed some of its principles, including releasing results as soon as possible and including information about how the models reached their conclusions. But not for all of them: Just 10 of the 719 manuscripts included releases of the modelās chain of thought. For papers that people donāt understand, the mathematicians suggested the proofs should be formalized ā but just 42% of the proofs released by OpenAI had not undergone this process. Ultimately, itās still not clear that OpenAI is taking āresponsibility for ensuring that human understanding will followā when releasing its proofs, in accordance to the AGMAI principles. AGMAI suggested that OpenAI should help fund the work of human mathematicians who will be required to make the labās solutions meaningful in any real way. āProblems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is āsolvedā, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field,ā Terence Tao, a prominent mathematician who has criticized OpenAIās approach,wroteon social media after the release. That problem is exemplified bya paperreleased this week by mathematicians at the University of Cambridge and Kingās College in London that questions the way frontier labs are approaching these challenges. When AI models solve mathematical problems, they first create a ānatural languageā explanation, then try to express that result in Lean, a programming language that in theory confirms the accuracy of the proof by compiling it as code. However, there may be problems with the way the models translate their natural language proofs into code; this paper documents at least two discrepancies between the natural language proof and the Lean code behind the solution OpenAI has offered to a problem derived from the Navier-Stokes equations that describe the complex behavior of fluids. These discrepancies donāt necessarily disprove either solution, but they do raise questions on whether we can simply rely on models to formalize their own solutions without human involvement. Thatās one reason that AGMAI asked OpenAI to āinclude machine-readable metadata correlating the natural language and formal artifacts,ā something that the frontier lab did not do with these releases. āBecause of the phenomenon of mistranslations ā as highlighted in this paper ā the NL proof by OpenAI andother autoformalised Lean proofs should not prima facie be trusted without the same peer review process andscrutiny that other proofs are subjected to,ā the authors of the ālost in translationā paper conclude. Mathematicians stress that when new results are discovered by humans, they take responsibility for them and engage with the broader community through papers, talks, and seminars. That process increases understanding of the solutions, finds strategies that can be used to solve other problems, and allows the new knowledge to be applied in practical fields. When a model is prompted to solve a hard problem and spits out a solution, āthere is not human understanding of them at the point of release, and now the work begins,ā Harvard University mathematics professor Melanie Wood told TechCrunch.
read more



