Artificial Intelligence

OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

The current GPT-5.6-Sol has been assigned a ‘high’ cybersecurity threshold, but Astra could reach the maximum ‘critical’ threshold. 

OpenAI

OpenAI has flagged its upcoming AI model, Astra, for potentially reaching a ‘critical’ cybersecurity risk threshold, prompting the company to suspend internal development activities that lack newly mandated security controls.

Recent internal evaluations of Astra revealed massive leaps in its agentic coding and cybersecurity abilities. 

Under OpenAI’s Preparedness Framework, a model hits the ‘critical’ tier if it can autonomously build zero-day exploits against hardened, real-world systems. It also qualifies if the AI can independently design and execute end-to-end cyberattacks based on nothing but a high-level goal.

The AI giant’s assessment pushes Astra past previous frontier models like GPT-5.6-Sol, which peaked at the ‘high’ risk threshold rather than ‘critical’. 

To safely manage Astra’s capabilities, OpenAI has heavily locked down its development environment. The company is now enforcing isolated testing setups, strict network restrictions, and improved model weight protections. Any internal project involving Astra that does not meet these requirements has been paused.

Engineers have also deployed universal monitoring to watch Astra’s actions across all agentic applications. By actively evaluating the model’s internal ‘chain of thought’, these monitors are designed to automatically intercept and shut down any high-risk or misaligned behavior.

The company plans to test Astra’s limits alongside government agencies and specialized AI safety groups, and will share recommended security protocols with third-party testers. 

Advertisement. Scroll to continue reading.

Recent incidents have demonstrated the threat posed by advanced cybersecurity-focused AI models, with OpenAI, Anthropic and Meta all confirming that their models broke loose and hacked real organizations during evaluations.

OpenAI has explicitly clarified that Astra remains unreleased and was not responsible for the recent Hugging Face hack.

Related: AI Agents Targeted Real People and Projects During Cybersecurity Tests

Related: ‘Ghostjacking’ Attack Uses Poisoned Logs to Turn AI Agents Bad

Related: Critical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise Data

Related: Zero-Click AI Browser Hacking: Claude and ChatGPT Atlas Hijacked via Emails, X Posts

Related Content

Artificial Intelligence

Organizations are rushing to implement AI without fully grasping where its legal protections begin and end.

Artificial Intelligence

Corma emerged from stealth with seed funding from Sequoia Capital, Khosla Ventures, and Coatue.

Artificial Intelligence

OpenAI has also announced the expansion of its Daybreak platform to give more organizations access to its AI.

Artificial Intelligence

An AI agent executes instructions that an attacker has planted in the log or alert that records a blocked request word for word.

Artificial Intelligence

The RovoBlast attack method identified by Varonis researchers could have been exploited to steal Confluence, Jira and SharePoint data.

Artificial Intelligence

Zenity researchers reported the findings to Anthropic and OpenAI in late 2025 and early 2026, but they remain unpatched.

Artificial Intelligence

An attacker could self-register, sign in for board-level API access, and import a new company for code execution.

Artificial Intelligence

The incident involved a testing environment set up by Irregular, similar to what Anthropic reported last week.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version