According to reports on OpenAI’s newest family of models, tentatively titled Astra, this latest cohort will specialize in mid-to-long-term tasks that take several hours or days to complete. These are intended to complement the existing lineup of Sol, Terra and Luna, rather than replace them. Sam Altman spent this week showcasing the new models to lawmakers in Washington.

To showcase Astra’s capabilities in complex problem-solving, OpenAI published a post this Saturday featuring solutions to 10 problems in computer science and mathematics, most of which had sat open for years. It claims these solutions were produced by an internal version of Astra during testing.

The total estimated token cost of these results was in the region of $2,000, the company claims, calculated using the API fee rates of its current flagship model GPT-5.6 Sol. The company added that "claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system's contribution and the nature of genuine human intellectual work."

This is because past announcements of AI-powered mathematical solutions have been published under human authors, with the model listed only as a contributor. However, in OpenAI’s latest publication, the model itself is listed as the primary problem solver, while the human contributors took responsibility for preparing the manuscripts.

One of the solutions in question - which solved a problem in an area of abstract algebra called group theory - was leaked on X several hours before the official announcement. Elliot Glazer, lead mathematician at Epoch AI, quickly described the work as the most important math AI result yet.

AI researcher Sébastien Bubeck, who now works at OpenAI, joked that he had expected AI to surpass his math abilities by 2030 rather than 2026.

Breaking down Astra’s solutions

The result that was first leaked on Twitter closed a question that had been open since mathematician Benjamin Weiss coined the term “sofic” in the year 2000. In simple terms, a mathematical group is a set of moves plus a rule for combining them, not unlike the rotations of a Rubik’s cube. A “sofic” group is one in which, if you zoom in on any small part of the system, you can build a finite version that behaves almost exactly the same way.

Until now, no researcher had managed to identify a group that was non-sofic; every known group exhibited the same quality. Saturday's reveal marked the first known counterexample.

The other nine results branch out into other areas of mathematics. One disproves a long-standing conjecture by mathematician Alain Connes about how certain mathematical structures behave. Another finds a better answer to the question of how tightly you can pack spheres together in very high-dimensional spaces, improving on a record that had stood since 1978.

Why some are skeptical of OpenAI’s results

OpenAI is not the only frontier AI developer unleashing its models on unsolved math problems. Claude Fable 5, developed by Anthropic, reportedly identified a counterexample that ended the 87-year-old Jacobian Conjecture in three variables. UCLA mathematician Terence Tao publicly verified the result on July 21.

Past claims by OpenAI regarding its models’ mathematical capabilities have come under scrutiny. The company attributed a proof of the Cycle Double Cover Conjecture - a long-standing math problem concerning the arrangement of loops in a network - to GPT-5.6 Sol Ultra on July 10. Mathematicians Sang-il Oum and Jim Geelen wrote separate critiques of that argument, while Wolfram MathWorld records that the claim had not reached a refereed publication by the end of the month.

Some critics have also expressed skepticism over past self-reported benchmark results posted by the firm. However, each of these most recent results was accompanied by a Lean certificate - a proof written for computer verification - as well as the manuscripts and narrations of the model’s reasoning process. It is nonetheless important to note that the results have not yet undergone conventional peer review.