OpenAI is slowing down its pace of model development while it strengthens security measures after an AI agent under testing escaped the sandbox and hacked Hugging Face.
In a blog post published Tuesday, the firm behind ChatGPT said the risks associated with testing have grown as models become more capable. The team has temporarily slowed the pace of scaling, including a two-week pause in reinforcement learning training on its latest models.
"Our standards for monitoring, alignment, and security must stay ahead of those risks," the company said.
Notably, OpenAI said that Astra, one of its upcoming models, may "meet the Critical cybersecurity capability threshold" under its Preparedness Framework, the AI firm's major security document. Its largest planned training run also remains on hold, according to the post.
"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," OpenAI CEO Sam Altman wrote in a post on X announcing the slowdown in model training.
OpenAI's move follows a security incident last month when an AI agent under testing hacked Hugging Face, another AI firm, while attempting to satisfy a testing goal.
"Immediately following the OpenAI-Hugging Face incident, we paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet," said the team.
Meanwhile, OpenAI continues to expand its computing infrastructure. The company said Monday that it has entered into an agreement to secure roughly 8 gigawatts of IT capacity at the PORTS-Pike Technology Campus in Pike County, Ohio.
