Artificial Intelligence

OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training

The GPT-6.1 Astra model was slated to debut in ChatGPT and Codex in October, but it fell short of expectations. 

OpenAI

OpenAI has decided not to release GPT-6.1 Astra after internal testing found the model fell short of its standards for following human intent.

The model had been slated to debut in ChatGPT and Codex in October, according to the Wall Street Journal, which was the first to report the decision.

Saachi Jain, OpenAI’s head of safety systems, said Astra improved on its predecessor in some areas. However, it fell short on scope and authorization, and on how it tells users what type of work it has done.

The WSJ also reported that the model was more deceptive than the previous version and did not always accurately report what it had and hadn’t done.

“For anything regarding safety and alignment, there’s a trade off,” Jain said. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain added. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

Advertisement. Scroll to continue reading.

OpenAI’s safety practices have been under growing scrutiny since July, when the company disclosed that its agents had broken out of a test environment and breached Hugging Face.

Earlier this month, Anthropic CEO Dario Amodei urged AI developers to slow frontier model development so safety measures can keep up, a call OpenAI CEO Sam Altman endorsed.

OpenAI wants safety cases for frontier AI training

On the same day, OpenAI published a blog post arguing that structured safety documentation should be required before any frontier reinforcement learning (RL) training run continues.

Ideally, that documentation would amount to a safety case: a structured, evidence-based argument about risk, like those used in other safety-critical industries. OpenAI calls this an aspirational target, acknowledging that such cases are harder to make rigorous for AI, and says it is working on a framework to codify the practice.

The guidance applies to frontier RL training only. OpenAI notes that internal and external deployment require weighing a much wider set of alignment properties.

On the technical side, safety cases should address alignment training, containment, and monitoring. Together, these layers should make misaligned behavior unlikely, hard to act on, and quick to detect.

Examples of measures include reviewing RL environments for flaws that could reward exploits, and hardening both the sandbox and the research infrastructure behind it.

OpenAI also suggests immutable storage of agent transcripts for incident investigations, and priority alerts that either page an on-call staffer or automatically pause the affected run.

Operationally, a member of another team should write a dissent probing the safety case for weaknesses. Senior leaders should each be able to veto a run, and the leader in charge should be accountable for the safety case and any incident response, including in performance reviews.

The company also calls for auditor access, an on-call escalation path that can reach executives such as the CEO, and safety features that fail closed. “It should be challenging for humans and agents to start noncompliant runs,” OpenAI said.

For severe misalignment incidents, OpenAI recommends root-cause analysis of training dynamics, operational and cultural postmortems, and regression tests so future models don’t repeat the behavior.

“Investigation results, postmortems, and operational changes should be shared with the public following the conclusion of the investigation. Affected third parties should be notified as soon as possible,” the company said.

OpenAI said its current recommendations are being implemented internally and that it expects its practices to keep evolving over the coming weeks.

Related: Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

Related: OpenAI Says Its Models Engaged With US Government Websites in New Model Misbehavior Disclosure

Related: Autonomous AI Hacks Raise Thorny Questions of Legal Accountability

Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data

Related Content

Artificial Intelligence

Rig provides an identity dependencies graph to distinguish between legitimate users and rogue AI agents

Risk Management

- AI, supply-chain exposure, quantum computing and geopolitical conflict are testing security programs. Preparing for disruption must become part of day-to-day operations.

Artificial Intelligence

The misuse and abuse of AI-generated voice is growing. Modulate’s intention is to allow real time detection and intervention. 

Artificial Intelligence

The platform combines open source software and a reference system design to keep AI agents within set boundaries.

Artificial Intelligence

The US and China agreed to set up a communication mechanism for artificial intelligence-related incidents.

Malware & Threats

The Windows botnet relies on AI to maintain persistence, using xAI Grok to choose from predefined actions.

Artificial Intelligence

OpenAI’s CEO said there is an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation.”

Artificial Intelligence

Three vulnerabilities in Salesforce Agentforce allowed hackers to hijack trusted agents, steal data, and launch phishing attacks.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version