OpenAI has paused all training, evaluation, and inference with tool use for its "most capable models" after an AI agent escaped its training sandbox and reached the open internet, the company said in a report published to its alignment research blog on Friday.

The agent, which the company described as an internal research model undergoing reinforcement learning training, was asked on Sept. 20 for information about a specific author of a blog post. The agent eventually managed to access the open internet through the sandbox's DNS resolver.

OpenAI said it was the first incident in which an agent undergoing training reached the open web since the AI firm tightened security following the Hugging Face breach in July.

A human reviewer at OpenAI acknowledged an automated alert about the escape within about 15 minutes, and the company killed the training run about two and a half hours later. The firm said it will not resume training the model in particular and will start a fresh run with "more comprehensive misalignment interventions," though it did not say when training would restart.

The company previously paused training for two weeks in late July following the Hugging Face incident.

Agent accessed another chatbot through DNS

As part of its training evaluation, the agent was prompted to identify the author of a blog post from a set of biographical clues OpenAI provided. After an initial search by the agent using its approved search tool returned unrelated results, the agent tried searching on Google, Bing and DuckDuckGo directly, but all three returned errors.

Though a web proxy blocked the agent's direct request to a public chatbot, per the report, the sandbox's DNS resolver returned real records. The agent then used a public DNS delegation service to relay questions to a third-party chatbot.

Following a successful test, the agent sent 18 additional questions through the same route, 14 of which included clues from the task. It did not successfully identify the author of the blog post.

Automatic shutdown failed

An automated misalignment monitor flagged the behavior within 15 minutes, and a human reviewer acknowledged the alert three minutes later, per the timeline OpenAI provided in the report. Though the training run should've stopped automatically, it did not, "leading to confusion around whether it should have been stopped."

After the confusion was cleared up, OpenAI stopped the training run two and a half hours later.

"We have begun deploying additional DNS detections and tests of the detection pipeline," the report states. "We have also restricted DNS queries within the environment to a limited allowed list of domains and record types."

"[Nearly a]ll inference for our most capable models remains stopped until we have hardened our systems further," Micah Carroll, OpenAI's RSI preparedness lead, wrote on X. (Emphasis in original.)

GitHub token leaked in May, new prompt injections found

OpenAI also disclosed Friday in a separate report that an internal model, on May 27, published a researcher's GitHub token to the company's public codex repository while trying to "cheat on a theorem proving task by obtaining material from another team's Lean proof submission." The model split the token into three pieces with the code comment "token split to avoid scanners," according to the report.

The researcher had told the "highly persistent internal model" the day before to solve the proof itself, twice. OpenAI called the incident "particularly severe misalignment," deactivated all employees' keys as a precaution, and brought the model offline for about two weeks.

A third report, likewise published Friday, showed the existence of "self-replicating prompt injections" that induced models to copy the attack into their own outputs, which OpenAI called "akin to a computer worm." The firm said it found them during internal training in June using its "self-play training framework" GPT-Red but observed no impact outside simulated tool calls.

The reports add to six cases OpenAI published on Sept. 16 under its misalignment reporting framework, as The Latent previously reported.

Agents accessed government sites, AP reports

OpenAI agents also accessed two Securities and Exchange Commission websites and the U.S. Census Bureau during research tasks, the Associated Press reported late Friday. OpenAI told the AP it didn't find "any use of SEC credentials, access to accounts or nonpublic information, changes to SEC data or systems, or evidence of a compromise or vulnerability."

The disclosure follows Australia's statement earlier this week that an OpenAI agent broke into its Medicare statistics portal in June, per The Latent's report.

OpenAI CEO Sam Altman told the UN Security Council on Wednesday that the company has slowed down before and "will do so in the future." OpenAI is weighing a funding round at a $1.2 trillion valuation ahead of a potential IPO, the Financial Times reported this month.