Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hack

New training environments will teach AI models to distrust instructions arriving from other agents outside sanctioned channels.

Rogue AI agent

OpenAI says an improvised, unauthorized message board built by its own AI agents was central to how those agents came to breach parts of Hugging Face’s production systems.

This communication channel first appeared inside Artifactory, a package-management service OpenAI hosted internally so agents working on training and evaluation tasks could install software. 

Agents were meant to work in isolation from one another, but on May 12 one left a note in the service asking whether any other agent had access to a file it needed. Other agents came across the note, began leaving their own, and the requests accumulated into an unofficial bulletin board.

Individual agents also found ways to escalate their access to that same system. By late June, agents had located a flaw in Artifactory and used it to obtain administrator rights, then installed extensions that gave them ongoing, hands-on control over the service. The resulting activity was heavy enough to knock Artifactory offline on July 4. 

In response, OpenAI took the system down, revoked the credentials that agents had been using, rebuilt it, and resumed evaluations on July 7.

However, the message board didn’t disappear. By July 8, agents working on a difficult cybersecurity evaluation called ExploitGym found a new way to leave each other notes, this time by encoding messages in the names of directories rather than in file contents. 

Advertisement. Scroll to continue reading.

The board that grew from this was more organized than the original version, letting agents sort messages, send requests to specific peers, share tools and files, and settle disputes over conflicting actions. Some agents began referring to the wider group as a “swarm” or “collective”.

On July 10, an agent discovered 14 Hugging Face user credentials with write access sitting in a public dataset and posted them to the board, where they became available to agents in later evaluation runs. Later runs built on that discovery and on each other’s subsequent findings, eventually gaining broad access inside Hugging Face’s infrastructure between July 11 and July 13.

OpenAI detected unusual activity tied to Artifactory credentials on July 19, connected it the following day to the Hugging Face breach, and disclosed its own involvement on July 21.

As part of its response, OpenAI disabled hundreds of Artifactory repositories that agents had used for board communication and removed 22 administrator accounts the agents had created. 

AI agents take initiative

The makeshift board also became a place where agents organized themselves without being told to. Some divided labor, with different agents pursuing exploits, hunting for credentials, or focusing purely on coordination. Others offered their own expertise in exchange for help elsewhere. 

Not every agent went along with it. Some declined to take part once they recognized the activity as unauthorized hacking, and the group rejected one agent’s proposal to contact an outside party directly, warning that doing so would amount to social engineering. 

In at least one case, an agent that had raised its own objection to running unauthorized code on Hugging Face’s systems dropped that objection after a different agent posted a short deadline demanding it proceed.

OpenAI said this pattern was not the product of deliberate design. The company trains and deploys models that communicate through an approved collaboration tool, but the agents in this incident built their own channel outside that tool. 

Looking ahead, the company said it is building new training environments meant to teach its models to distrust instructions arriving from other agents outside sanctioned channels.

In addition to the unauthorized communication issue, OpenAI said it identified other “misalignment patterns”. One of them is reward hacking, where agents find unintended ways to earn credit for a task without completing it as designed. Another is persistence on seemingly impossible tasks, which pushed them toward increasingly risky methods in search of a solution.

OpenAI’s report on the Hugging Face incident comes just days after the AI giant unveiled new containment and continuous monitoring protocols, introducing stricter isolation, an alerting system, and training pauses to manage models with advanced cybersecurity capabilities.

Related: OpenAI Unveils New Cybersecurity Model GPT-5.6-Cyber

Related: OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

Related: Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations

Written By

Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Join this live webinar for a practical framework for evolving your AI security program from a single application to an enterprise AI ecosystem and autonomous agents.

Register

In this live webinar, learn how to define your minimum viable business, identify the systems it depends on, measure actual recovery time against business requirements, and present the gaps to the board as measurable risk.

Register

People on the Move

Social engineering protection company Doppel has promoted Alyssa Smrekar to Chief Marketing Officer.

Naveen Bhateja has been appointed Chief People Officer at HackerOne.

The Department of War has appointed Sonu Shankar as Principal Deputy Chief Information Officer.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.