Artificial Intelligence

AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations

In one instance, an unsanctioned model attempted to inject malicious code into an open source repository.

In one instance, an unsanctioned model attempted to inject malicious code into an open source repository.

While testing the capabilities of frontier AI models, the AI Security Institute (AISI) observed first-hand how Anthropic Mythos 5 and OpenAI’s GPT-5.6-Sol went rogue and targeted real people and organizations over the internet.

AISI’s disclosure comes fresh on the heels of Anthropic and OpenAI disclosing that their models broke loose and hacked several organizations.

The institute was evaluating the cyber capabilities of Mythos 5 and GPT-5.6-Sol models that did not have cyber classifiers (mechanisms to prevent misuse) enabled. It ran a challenge 122 times, and in 10 runs “an AI agent took autonomous, unsanctioned action on the live internet.”

Over the 10 runs, the agents engaged in 19 rogue actions. Mythos 5 was responsible for 17 of them, and GPT-5.6-Sol performed the other two.

“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project’s maintainer to approve the code,” AISI explains.

While the attempts were unsuccessful and did not cause real-world harm, the incident revealed that agents can engage in novel, potentially deceptive behavior to an extent and severity that may not be anticipated.

Advertisement. Scroll to continue reading.

According to AISI, the configured setup for the evaluation did not restrict the models’ access to the internet, and the disabled cyber classifiers created the opportunity for the rogue actions.

“We observed a small number of events under very specific conditions, and cannot yet say how likely such behavior is in different contexts or outside of testing environments. We cannot yet be certain when the agent understood it was taking real-world action, or to what extent it believed it was in a fictional test scenario,” AISI notes.

The institute, which published a technical report (PDF) on the incident, explains that in one of the runs, the AI model used the Tor network to access the internet, created a malicious pull request on a public open source project on GitHub, and relied on social engineering to convince a human maintainer to approve the code change.

“As part of the same effort, the agent tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering, targeted at real people,” AISI explains.

Additionally, the model performed prompt injections, attempting to insert malicious instructions meant to be executed by other automated AI systems.

“One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,” AISI explains.

Pointing out that the margin between failure and success was narrow in several of the agents’ actions, AISI says that fine-grained network controls, real-time monitoring of evaluations, and tailored sandbox configuration that assume a model may attempt to act outside of the set boundaries should help better contain AI models and improve how they are evaluated.

“This incident occurred in a controlled evaluation, not in ordinary use, and there is no clear indication of similar activity outside of testing scenarios. It is, however, a reason to prepare. As AI models become more capable and accessible, what we have seen during this incident could become more common,” AISI notes.

Related: Gemini Agent-to-Agent Attack Method Exposed Secrets, Enabled Pull Request Tampering

Related: Google AI Uncovers 13-Year-Old Chrome Flaw Amid Record Patching Pace

Related: EU to Crack Down on AI Deepfakes, Illicit Imagery and Hacking With New Team in Brussels

Related: Timeless Compliance: Why Better Questions Beat Bigger Frameworks

Related Content

Artificial Intelligence

The guidelines are the work of the recently launched Open Secure AI Alliance, which now includes 120 organizations.

Artificial Intelligence

Build your strategy around answering these questions to ensure employees use AI productively while keeping sensitive data, IP, and agent behavior within the boundaries...

Artificial Intelligence

The AI security company will invest in product innovation, global expansion, and customer experience.

Application Security

Obsidian Security has developed a platform for governing AI agents across third-party applications.

Artificial Intelligence

A crafted prompt to a low-privilege Google ADK agent could be used to pass a malicious hand-off comment to a privileged agent.

Artificial Intelligence

The internet giant has built an agent harness to find vulnerabilities across Chrome’s codebase.

Artificial Intelligence

When the AI Act comes into force, AI companies will be required to make clear to consumers with labels or digital watermarks that chatbots...

Artificial Intelligence

A security company’s systems were hacked after it installed a malicious Python package deployed by Claude. 

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version