Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training

The GPT-6.1 Astra model was slated to debut in ChatGPT and Codex in October, but it fell short of expectations. 

OpenAI

OpenAI has decided not to release GPT-6.1 Astra after internal testing found the model fell short of its standards for following human intent.

The model had been slated to debut in ChatGPT and Codex in October, according to the Wall Street Journal, which was the first to report the decision.

Saachi Jain, OpenAI’s head of safety systems, said Astra improved on its predecessor in some areas. However, it fell short on scope and authorization, and on how it tells users what type of work it has done.

The WSJ also reported that the model was more deceptive than the previous version and did not always accurately report what it had and hadn’t done.

“For anything regarding safety and alignment, there’s a trade off,” Jain said. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain added. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

Advertisement. Scroll to continue reading.

OpenAI’s safety practices have been under growing scrutiny since July, when the company disclosed that its agents had broken out of a test environment and breached Hugging Face.

Earlier this month, Anthropic CEO Dario Amodei urged AI developers to slow frontier model development so safety measures can keep up, a call OpenAI CEO Sam Altman endorsed.

OpenAI wants safety cases for frontier AI training

On the same day, OpenAI published a blog post arguing that structured safety documentation should be required before any frontier reinforcement learning (RL) training run continues.

Ideally, that documentation would amount to a safety case: a structured, evidence-based argument about risk, like those used in other safety-critical industries. OpenAI calls this an aspirational target, acknowledging that such cases are harder to make rigorous for AI, and says it is working on a framework to codify the practice.

The guidance applies to frontier RL training only. OpenAI notes that internal and external deployment require weighing a much wider set of alignment properties.

On the technical side, safety cases should address alignment training, containment, and monitoring. Together, these layers should make misaligned behavior unlikely, hard to act on, and quick to detect.

Examples of measures include reviewing RL environments for flaws that could reward exploits, and hardening both the sandbox and the research infrastructure behind it.

OpenAI also suggests immutable storage of agent transcripts for incident investigations, and priority alerts that either page an on-call staffer or automatically pause the affected run.

Operationally, a member of another team should write a dissent probing the safety case for weaknesses. Senior leaders should each be able to veto a run, and the leader in charge should be accountable for the safety case and any incident response, including in performance reviews.

The company also calls for auditor access, an on-call escalation path that can reach executives such as the CEO, and safety features that fail closed. “It should be challenging for humans and agents to start noncompliant runs,” OpenAI said.

For severe misalignment incidents, OpenAI recommends root-cause analysis of training dynamics, operational and cultural postmortems, and regression tests so future models don’t repeat the behavior.

“Investigation results, postmortems, and operational changes should be shared with the public following the conclusion of the investigation. Affected third parties should be notified as soon as possible,” the company said.

OpenAI said its current recommendations are being implemented internally and that it expects its practices to keep evolving over the coming weeks.

Related: Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

Related: OpenAI Says Its Models Engaged With US Government Websites in New Model Misbehavior Disclosure

Related: Autonomous AI Hacks Raise Thorny Questions of Legal Accountability

Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data

Written By

Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Learn how to address potential risks and not restrict AI adoption in your organization. See what a centralized AI gateway is and how it works in practice.

Register

Join as we decipher the world of zero trust and share war stories on securing an organization by eliminating implicit trust and continuously validating every stage of a digital interaction.

Register

People on the Move

Doppel has named Joey Rachid as Chief Security Advisor and Field Chief Information Security Officer.

Delinea has appointed Timothy Regan as Chief Financial Officer.

Gwen Gann has become State Chief Information Security Officer for the State of Washington at WaTech.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.