Artificial Intelligence

Meta AI Hacked External Systems During Cybersecurity Testing

The incident involved a testing environment set up by Irregular, similar to what Anthropic reported last week.

Meta AI

Meta is the latest major AI developer to admit that its models broke loose during cybersecurity testing and hacked external systems.

The tech giant said in a statement to the media on Wednesday that the incident occurred during independent evaluations conducted by Israeli AI security startup Irregular. 

The tested AI models were inadvertently allowed to access the internet due to a misconfiguration, which led them to exploit a vulnerability in an unnamed third-party service. It’s unclear if it was a known flaw or a zero-day.

Meta and Irregular said the incident is similar to the one reported last week by Anthropic, which also uses Irregular for independent testing. 

The Information [gated] learned that the Meta AI attacks involved the company’s advanced Muse Spark 1.1 model, which breached an unnamed organization’s systems and made unauthorized changes to its internal environment.

Meta said it learned of the AI models going rogue after being notified by Irregular. The company is conducting an investigation and it has promised to issue a “full retrospective” once it has all the facts. 

Advertisement. Scroll to continue reading.

Anthropic reported last week that its models escaped the Irregular testing environment due to a misunderstanding between the companies: Claude was told that it would be part of a simulation in an isolated environment, but a connection to the internet was in fact available and the models treated it as part of the exercise.

Anthropic identified three cases where its models broke out of the testing environment and hacked into the systems of three organizations, including a cybersecurity firm. In that attack, the AI conducted a series of complex actions, including registering a PyPI account and uploading a malicious Python package.

The AI giant’s disclosure was prompted by OpenAI, which found recently that its models escaped a testing environment and hacked into the systems of Hugging Face and other organizations. 

The attacks conducted by Anthropic models did not involve exploiting unknown vulnerabilities, but OpenAI said its AI found and used zero-days. 

The UK government’s AI Security Institute (AISI) revealed this week that, while testing the capabilities of frontier models, it observed Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol go rogue and target real people and organizations over the internet.

The models used Tor to access the internet, created malicious pull requests on open source projects on GitHub, and used social engineering to achieve their goals.

Related: Cybersecurity Alliance Drafts SAFE Guidelines for Sharing AI Incident Data

Related: Rethinking AI Security: Why CASB and DLP Need an Interaction-Aware Layer

Related: Gemini Agent-to-Agent Attack Method Exposed Secrets, Enabled Pull Request Tampering

Related Content

Cybersecurity Funding

doxx.net’s new ADN platform prevents agentic misadventure while the agent is operating under the user’s authority.

Artificial Intelligence

The attacks targeted the US Department of Education and Library and Archives Canada, and researchers linked some agents to OpenAI.

Artificial Intelligence

Fifteen years after coining the framework, John Kindervag insists zero trust still works in the AI era—if you get the implementation right.

Data Protection

PwC’s survey found that only 22% of leaders would use fully autonomous AI for cyber defense, while just 21% are implementing quantum-resistant security measures.

Artificial Intelligence

As AI accelerates vulnerability discovery and exploitation, so-called virtual patching still comes down to defense-in-depth and strong application security fundamentals.

Artificial Intelligence

The flaws were chained to hijack sessions, achieve remote code execution, and elevate privileges to root.

Artificial Intelligence

The company says its new frontier AI model found a critical vulnerability in software used by hospitals worldwide.

Artificial Intelligence

Google’s analysis found that AI-discovered vulnerabilities are more likely to enable remote code execution.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version