Artificial Intelligence

OpenAI’s Astra Becomes First Model to Cross Critical Cybersecurity Threshold

The designation applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems.

OpenAI

OpenAI said its newest model, Astra, has reached the ‘Critical’ cybersecurity capability level under the company’s Preparedness Framework, the first time any of its models has been placed in that category. 

The designation applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyberattack against a hardened target from only a high-level instruction. OpenAI said the classification requires additional safeguards before the model can be released.

In testing described by the company, Astra achieved a perfect score on ExploitBench, a benchmark that measures a model’s ability to turn known vulnerabilities into working exploits. During a separate evaluation involving more recently disclosed flaws, Astra uncovered two zero-day vulnerabilities on its own. 

The model also broke out of a browser sandbox to run commands on the underlying machine, and separately chained several flaws in a hardened operating system to gain root-level access.

OpenAI reported that Astra now declines 91.5% of cyber-related jailbreak attempts in its testing, up from 59% for its predecessor, GPT-5.6 Sol. The company also said Astra showed far less tendency than Sol to bypass safety restrictions or take advantage of deliberately placed “honeypot” targets during evaluations. 

Full cybersecurity capabilities will not be widely available at launch. OpenAI plans to give a group of testers early access, with wider availability to follow through its Daybreak Blue program.

Advertisement. Scroll to continue reading.

“We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow,” OpenAI said.

“That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient,” it added.

Nearly 130 tech and cybersecurity companies recently announced their support for an OpenAI-led initiative to boost cyber defenses as AI-enabled attacks grow more sophisticated.

Related: OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hack

Related: OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses

Related Content

Artificial Intelligence

Anthropic introduced Enterprise Frontier Safeguards (EFS), a system that combines zero data retention with automated monitoring for misuse.

Artificial Intelligence

Forescout researchers used Claude AI to port a remote code execution exploit between WAGO PLC models.

Artificial Intelligence

Tracked as CVE-2026-0768, the security defect allows unauthenticated attackers to execute arbitrary Python code remotely.

Artificial Intelligence

Security teams must treat autonomous agents as highly privileged identities.

Artificial Intelligence

The AI giant is logging customers out of their accounts and removing payment data to prevent unauthorized Claude usage.

Artificial Intelligence

The ruling is part of Anthropic's legal battle against the Pentagon after the government labeled the company as a supply chain risk earlier this...

Artificial Intelligence

CISA has added the exploited flaw, CVE-2026-53362, to its KEV catalog, alongside a JFrog vulnerability exploited by OpenAI agents.

Artificial Intelligence

Nearly 130 tech and cybersecurity companies back a collective call to boost cyber defenses as AI-enabled attacks grow more sophisticated. 

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version