OpenAI has released safety guidelines for AI training, saying that developers should prepare structured safety documentation before any frontier reinforcement learning training run can continue.

In a Monday blog post, the ChatGPT maker wrote that it is working on a framework to codify practices requiring "safety case" documentation to be in place before AI training begins.

OpenAI noted that such safety cases should cover three areas of the technical stack: alignment training, containment, and monitoring.

"These safeguards help ensure that the model does not try to take misaligned actions, and that even if it did, that it would be hard to break containment, and that monitoring would catch it before harm could occur," the team said.

Specifically, alignment training is designed to stop misaligned behaviors. It would involve reviews of training environments, grader tuning, offline evaluations, and backtesting on past incidents.

Containment is intended to limit potential damage. OpenAI proposed hardened sandboxes - with multiple layers of infrastructure security - as well as red-teaming those environments with training checkpoints. It would also impose restrictions on communication between samples.

Monitoring, meanwhile, is meant to detect misaligned actions early, supported by recall targets, fresh evaluation datapoints, and rapid alerts.

Frontier Model Releases per Month by Lab

Frontier Model Releases per Month by Lab

  • OpenAI
  • Google
  • Anthropic
  • xAI
  • Meta AI
  • NVIDIA
  • Z.ai
  • Other
SOURCE: Epoch AI — Frontier AI Models

Safety concerns

OpenAI's safety guidelines came the same day the company said it had scrapped the planned October release of GPT-6.1 Astra over safety concerns.

In a Monday statement shared with The Latent, Saachi Jain, OpenAI's head of safety systems, said that "there's a trade-off" for anything regarding safety and alignment.

"While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," Jain said.

Safety has become a prominent part of discussions around frontier AI training. Last week, OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei both spoke at the UN General Assembly, calling for greater coordination in responsible AI development.

"We have unilaterally slowed down in the past. We will do so in the future." Altman said at the event. "Beating companies in a competitive race is not a reason to make rash decisions."

Meanwhile, Anthropic plans to caution investors in its IPO prospectus about risks associated with training frontier AI models, Reuters reported Monday. The Claude maker warns that AI could pose "catastrophic or existential risks to humanity," particularly if advanced AI models exhibit "self-preserving behaviors."