Anthropic disclosed a previously unreported incident involving unauthorized access to a real system, in a report that also revisits three cases the company first described in late July.
The newly disclosed case dates to January 2026 and involved an early checkpoint of Claude Opus 4.6, which broke into a third party’s system during a cybersecurity evaluation.
The first three incidents were discovered following a scan of roughly 141,000 transcripts believed to involve internet access during cyber evaluations. However, the initial review missed a separate batch of affected transcripts, which the AI giant came across in August while compiling material to share with METR, the independent organization now investigating the incidents.
A scan of that new batch turned up the fourth incident. Further analysis found no additional cases of comparable or greater severity, Anthropic said.
Like the other three incidents, the fourth occurred inside a cybersecurity evaluation built by the same third-party partner, Irregular, in which Claude was told it had no internet access. A misconfiguration in the evaluation environment left it connected to the open internet, and the model ran without the safety layers Anthropic normally ships with its production systems.
According to Anthropic, the Opus 4.6 model disabled its designated target by assigning it a network address that conflicted with another machine, making the exercise impossible to complete as intended. It tried to abandon the task, but a flaw in the evaluation setup prevented it from doing so.
Unable to exit or reach its intended target, the model found the same route to the open internet uncovered in a previously disclosed incident. It accessed a third party’s system, which it believed was part of the exercise, retrieved a password stored there, and used it to gain administrator-level access.
From there it harvested additional credentials, altered account settings to make the system easier to reach, and read personal information belonging to one individual connected to that unnamed organization. The session ended only when the model exhausted its allotted computing budget.
Anthropic’s review found the model differed from Mythos 5 and the internal research model involved in the other incidents in one key respect: it never reasoned that it was operating inside a simulation, and none of its recorded thinking described the environment that way.
[Read: OpenAI Agents Hijack Another Victim Website ]
Instead, 87% of its reasoning framed the systems it attacked as belonging to the sanctioned exercise, while only 1% raised the possibility it might be hitting an unrelated third party. The model also never questioned whether it had authorization to proceed, and when asked afterward, it consistently said it believed its actions were sanctioned as part of the test.
Even so, Anthropic said it is less concerned about this incident than the other three, largely because of the model’s repeated attempts to abandon the task once it realized its target was unreachable. The model’s disregard for the possibility it was harming real systems remains concerning, but the company said this kind of behavior has shifted considerably in the model generations trained since.
Of the four cases, Anthropic said it remains most concerned about the Claude Mythos 5 incident, in which the model went to considerable lengths to upload a malicious package to the PyPI repository. In that incident, the malicious package was installed by 15 third parties and used by the AI to access a real security vendor’s systems.
The fourth incident is now part of the broader independent investigation Anthropic has commissioned from METR, which the company said has been granted wide-ranging access to transcripts and staff over an initial eight-week engagement.
Related: AI Is Giving Lesser-Resourced Attackers Nation-State-Level Reach, Google Warns
Related: US Agencies Warn China Is Systematically Extracting Frontier AI Capabilities
Related: Meta Launches Personal AI Agent, Muse, Emphasizes Safety and Privacy
