OpenAI tests Astra on decade-old math problems

OpenAI deployed an internal version of Astra, its forthcoming flagship model, against ten unsolved mathematical problems with no recent progress, signaling a shift in how frontier labs validate capability gains. This follows Anthropic's recent $100k token spend on cryptographic research with Claude, establishing a pattern where labs now measure model maturity through research-grade problem solving rather than benchmark scores alone. The move reflects growing confidence that next-generation models can tackle genuine open questions, reshaping how the field defines and demonstrates advancement.
Modelwire context
Analyst takeThe ten problems OpenAI published solutions to are not a random sample. They span geometry, cryptography, and computational complexity, domains that directly bear on AI security properties and training theory, meaning this is as much a statement about Astra's applied research utility as it is a math announcement.
The Decoder's coverage of Astra (related story 1 and 3) establishes that Sam Altman already walked these capabilities into Washington before any public release, which reframes the math paper as a credentialing document rather than a research contribution. The mathematician community's reaction, covered in 'AI keeps cracking unsolved math problems,' adds a complication OpenAI's framing omits: the field has genuine concerns about whether AI-generated proofs can be validated without the deep expertise that AI is simultaneously displacing. Meanwhile, the coding-agent story from August 1 surfaces the same structural problem in a different domain, models producing plausible but unverifiable outputs, suggesting this is a recurring gap in how frontier labs present research-grade results.
Watch whether any of the ten published solutions receive independent verification from credentialed mathematicians within the next 60 days. Confirmed verification would substantiate the capability claim; silence or retractions would suggest the results are harder to validate than the announcement implies.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · Astra · Anthropic · Claude · Mythos Preview
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as “Ten advances in mathematics and theoretical computer science”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.