Sam Altman's OpenAI said its new Astra model is set to release "soon," following several weeks of safety-related delays, though access to its most advanced cybersecurity capabilities will be more limited.
In a blog post, the firm said that Astra has become the first model that meets the "critical cybersecurity capability threshold" under its Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
"Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use," OpenAI said.
In testing, Astra achieved a perfect score of 100% on ExploitBench, which evaluates the model's ability to develop exploits from known vulnerabilities. In a separate internal benchmark, it discovered and used two zero-day vulnerabilities as part of an exploit chain, according to the firm, with OpenAI in the process of disclosing those vulnerabilities to their maintainers.
Learning from Hugging Face
In July, Hugging Face encountered a security incident, with OpenAI disclosing that one of its AI agents under testing escaped its sandbox and hacked the platform to fulfill a task. Following that incident, OpenAI said last month that it is slowing down its pace of frontier model development while it strengthens security measures, including internal work on Astra.
While Astra was not involved in the Hugging Face hack, OpenAI said it has incorporated learnings from that incident into its safety approach.
"We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity,” it said.
For users, this means extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity, according to the firm.
"We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects," OpenAI said. "The models that follow Astra will demand more of us. We will take the time and do the work needed to meet that responsibility."
