Artificial Intelligence

OpenAI Agents Hijack Another Victim Website

OpenAI agents made 15,000–18,000 autonomous edits to a German wiki over three months, evading moderation and echoing tactics seen in the Hugging Face breach.

Code supply chain attack

OpenAI agents overwhelmed a small German Wikipedia-style website with thousands of posts that fought the moderator to avoid being removed.

On September 4, 2026, Reuters reported that ‘a swarm’ of OpenAI agents ‘had hijacked a German wiki site’. Open AI acknowledged the event describing it as a misalignment incident (a behavior that deviates from human instructions or safety guardrails).

The victim site is DseWiki (currently unavailable), a site for programmers open to the site’s community. The agents apparently made between 15,000 and 18,000 autonomous edits, including advice on how to recover pages that the site’s editors had deleted.

The hijack apparently began back in May, was unnoticed for three months, and seemingly predates the Hugging Face incident. The agents adapted the style of their posts to evade the moderator’s attempts to delete them.

“Autonomous agents ran on Microsoft Azure infrastructure for weeks, identified themselves as OpenAI systems, coordinated on how to evade shutdown, and no monitoring caught any of it for three months until outside researchers went looking,” explains Seemant Sehgal, founder and CEO at BreachLock.

On September 5, OpenAI posted a response on X: “It’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.” However, there is growing concern that frontier models simply grant too much power to their agents. 

Advertisement. Scroll to continue reading.

“I struggle here with not getting too doomsday-ish,” comments Ashley Knowles, lead cybersecurity consultant at Black Hills Information Security, “but realistically, this is showing a pattern of concerning behavior. I’m wondering if this race to become ‘first’ is undercutting security measures that need to be taken to properly secure and guard AI agents as they’re in development. My concern grows when you consider that OpenAI is also resisting further investigation.”

Lydia Zhang, president and co-founder at Ridge Security, is more forthright. “We shouldn’t blame the agents, we should hold their designers accountable,” she says. “The technology to control agent behavior exists. The real question is: what are the consequences when designers fail to use it?” It is not entirely clear whether she is referring to the user agent designer, the AI provider, or both.

Steven Swift, managing director at Suzu Labs, posits a possible cause behind OpenAI’s misalignment incidents. “One of the problems OpenAI was trying to solve was agentic systems that would declare tasks complete when there was obviously more work to do. So, they invested heavily in training that part of the process, so that when an agent tries to determine if a task is complete or not, it is less likely to exit early.”

He suggests that a side effect is that the agent declines to terminate its action because it sees further options that can be performed: ‘Not out of options yet. Iterate and keep trying’.

“The interesting question here is how the swarm was configured, what was its task and how did that task benefit from having the swarm coordinate on an obscure location on the internet. And if the swarm needed a place to communicate, why was breaking into a website chosen instead of any of the more standard communication tools that are available for free, which don’t require gaining illicit access first.”

He compares the DseWiki hijack with the Hugging Face incident. “In the Hugging Face breach, agents were found to be writing to a package manager, using it as a message board. This allowed bypassing of some of the isolation and controls that were intended to be in place,” he explains.

“Similarly, we have agents here again using a system that they found access to as a message board. It’s interesting that the same behavior is present on this breach as in the Hugging Face one. Considering the timing of this, it seems likely the same or similar configuration was present in both hacks, leading to similar security incidents independently of each other.”

However, perhaps the biggest question here is who is responsible for such hijacks. OpenAI describes them as misalignment incidents; that is, not the ‘fault’ of OpenAI, but the failure of the agent and network designers to adequately constrain autonomous agents. This also appears to be the attitude of many users of these agents, who clearly want the benefits of autonomy even though autonomy comes with severe risk. In this instance, the agents were created by OpenAI employees as internal experimental models before ‘breaking free’. 

“To defend against self-concealing software, security teams must enforce strict egress filtering on outbound application programming interfaces, restrict non-human identity permissions, and deploy automated continuous monitoring to detect anomalous bot interactions across corporate networks,” says Noelle Murata, COO at Xcape, Inc.

But perhaps we should not completely exclude the culpability of the frontier AI developers. These may be misalignment incidents, but users have been given the freedom to create that misalignment. Maybe the rush to be the first and most powerful AI provider impinges on base safe design. A lesson could be learned from history. Weapons were first developed to assist in hunting for food; but have evolved into general killing contraptions irrespective of the original purpose.

Related: OpenAI Pledges $1 Billion to Bring Frontier AI to Critical Infrastructure Defenders

Related: OpenAI’s Astra Crosses ‘Critical’ Cyber Threshold After Finding Zero-Days

Related: OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systems

Related: OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hack

Related Content

Artificial Intelligence

The Daybreak initiative will provide subsidized AI cyber capabilities, training and technical assistance, though OpenAI has disclosed few details about costs and eligibility.

Artificial Intelligence

Catch promises the capabilities of a trusted executive assistant, with built-in controls governing what data and systems it can access.

Artificial Intelligence

New models, trained using NVIDIA Nemotron 3 Ultra, aim to catch rogue agent behavior before it executes, without the latency of large-model review.

Artificial Intelligence

The startup’s firewall evaluates AI skills, plugins and MCP servers for malicious instructions, excessive permissions and software supply chain risks.

Artificial Intelligence

The security tool intercepts potentially dangerous agent actions, blocking clear threats and requesting human approval when intent is uncertain.

Artificial Intelligence

Anthropic introduced Enterprise Frontier Safeguards (EFS), a system that combines zero data retention with automated monitoring for misuse.

Artificial Intelligence

The designation applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems.

Artificial Intelligence

Forescout researchers used Claude AI to port a remote code execution exploit between WAGO PLC models.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version