Artificial Intelligence

Formula Predicts When AI Chatbots Are at Risk of Turning Bad

Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted.

AI hacking

Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted; and if predicted, prevented.

Since 50% of the world’s population carry devices that can run personal AI companions with no internet connection and limited security, they focused their research here. The lack of cloud-based safety filters, real-time telemetry, live monitoring, or the ability to patch weights once deployed provides a good test bed for analyzing AI’s chat-style transformer behavior when left to its own devices.

If there is a tipping point, it will stem from the AI’s Attention head. This is the computational component that determines which earlier AI tokens are most relevant when processing the current token. The decision lays the foundation for the next token and so on until the chatbot has finished responding to the user’s prompt. The token is loosely, but not precisely, related to individual words or parts of a word. Different tokens have different weights – an indication of the importance of individual tokens.

The primary argument is that the forward motion of tokens can slip from good to bad (that is, go rogue) due to competition in the Attention head between the conversation’s context and competing output basins. A conversation’s accumulated context can gradually shift Attention toward an undesirable basin until a tipping point is crossed and the model begins producing bad outputs.

Since the primary driver in a conversation is the user’s prompt sequence, it follows that both thoughtless and malicious prompts can hasten the AI’s slippage into bad outputs. This can be either immediate (one bad prompt) or delayed (the accumulated effect of poor prompts).

The researchers (Neil Johnson and Frank (Yingjie) Huo) have further developed a mathematical formula designed to estimate the tipping point represented by the number of good outputs that occur before the first undesirable output appears. Once that first bad output occurs, the AI is on the slippery slope to roguery since it starts to influence future tokens in unintended ways. Alignment with purpose is lost, and the AI can be categorized as ‘misaligned’.

The math formulaic tipping point was tested across seven open-weight transformer models built by three independent groups and ranging from 124 million (small model) to 12 billion (larger model) parameters. The results showed consistent alignment with predicted immediate versus delayed tipping regimes.

Advertisement. Scroll to continue reading.

The value of the research is that it provides an explicit mathematical explanation for an observable phenomenon. “The key insight is that the ‘Beast’ is as Simon worked out in Lord of the Flies, inside each AI already. And it can be triggered and amplified in a conversational setting with humans, or within a group of AIs themselves, producing a ‘Lord of the Fl(AI)es’ effect,” Johnson told SecurityWeek.

If triggered, he continued, “An AI-agent (‘Patient Zero’) tips to generating undesirable output, with no human required. Other AI-agents then receive that undesirable output. This tipping propagates among the AI agents who have contact with each other – like a spreading disease. But unlike a disease, there is no virus. It again needs no ‘bad actor’ human to place a virus, or to kick it off. No humans needed.” The result could be a single rogue agent or a swarm of rogue agents.

But he also proposes a solution: “A simple warning light placed within the AI before it produces its next output. This is easy for AI companies to insert. We have already inserted this in the open source models in our lab, but obviously we cannot get inside OpenAI and Anthropic’s closed AI models to do this.”

Without improved control and observation, AI apps can go rogue by both accident and malicious intent. Bri Frost, director of product management at Cloud Range, has separately explained the same phenomenon.

“When an AI agent hits a wall, the real question is whether it stops or starts improvising.” That wall can be caused by an unintended poor prompt or a bad actor’s intended malicious prompt injection. “An agent doesn’t need bad intent to create risk. It just needs a goal, access and no clear sense of where its boundaries are,” he comments. 

“That risk grows when the person giving instructions doesn’t know to set those boundaries. Every day, inexperienced users hand agents open-ended tasks without telling them when to ask questions, pause or get approval. Before giving an agent credentials or tools, teams should test it in a realistic environment, including with vague or poorly written prompts. Does it stay within its permissions? Does it try to work around restrictions? Does it escalate to a human when a task pulls it outside its lane? If you can’t answer those questions, the agent isn’t ready for that level of autonomy.”

The short answer is that it is almost impossible to see or prevent an AI going rogue. We may be able to reduce the incidence through extreme care, but we cannot guarantee it can be eliminated. But we do understand through the GWU research how and why it happens, and how we may provide early warning.

Related: Wikimedia Says Rogue OpenAI Agents Tried to Turn Its Tools Into Proxies

Related: Outerlimit Raises $16 Million to Stop Rogue AI Agents From Causing Harm

Related: Widened Scan Turns Up Fourth Rogue Claude Cyber Incident

Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data

Related Content

Artificial Intelligence

The cybersecurity startup will invest in product innovation, agentic research, and employee base expansion.

Artificial Intelligence

Anthropic is integrating the CVP and Project Glasswing into a single offering, with three levels of access to its most capable AI models.

Artificial Intelligence

Wikimedia looked into whether its own websites had seen activity like that disclosed by other organizations

Artificial Intelligence

Google has temporarily stopped accepting product vulnerability reports through its Open Source Software Vulnerability Reward Program (OSS VRP).

Artificial Intelligence

The announcement comes after Trump hosted top executives of AI companies at the White House last week.

Cybersecurity Funding

doxx.net’s new ADN platform prevents agentic misadventure while the agent is operating under the user’s authority.

Artificial Intelligence

The attacks targeted the US Department of Education and Library and Archives Canada, and researchers linked some agents to OpenAI.

Artificial Intelligence

Fifteen years after coining the framework, John Kindervag insists zero trust still works in the AI era—if you get the implementation right.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version