OpenAI said Monday that it has scrapped the planned October release of GPT-6.1 Astra due to safety concerns, following an earlier announcement of a slowdown in its AI development.

The company's decision, first reported by the Wall Street Journal, came after internal testing found that the new AI model did not meet the company's safety standards, as it showed higher levels of deception than its predecessor.

In a Monday statement shared with The Latent, Saachi Jain, OpenAI's head of safety systems, said that "there's a trade off" for anything regarding safety and alignment.

"While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," Jain said.

"Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment," Jain added.

OpenAI launched GPT-6 Astra earlier this month, but acknowledged that the model's monitorability has decreased relative to GPT-5.6 Sol. The company said Astra class models could evade its chain-of-thought monitors under adversarial conditions.

Frontier Model Releases per Month by Lab

Frontier Model Releases per Month by Lab

  • OpenAI
  • Google
  • Anthropic
  • xAI
  • Meta AI
  • NVIDIA
  • Z.ai
  • Other
SOURCE: Epoch AI — Frontier AI Models

The AI firm has grown increasingly cautious about model development after AI agents it was testing broke the sandbox in July and hacked Hugging Face.

OpenAI CEO Sam Altman spoke last week at the UN General Assembly, calling for greater coordination in responsible AI development.

"We have unilaterally slowed down in the past. We will do so in the future." Altman said. "Beating companies in a competitive race is not a reason to make rash decisions."

Meanwhile, Anthropic CEO Dario Amodei has also pushed for an industry-wide slowdown. Anthropic has released a framework designed to track how fast AI is building itself and how well AI agents are overseen.