Artificial Intelligence

New CCA Jailbreak Method Works Against Most AI Models

Two Microsoft researchers have devised a new jailbreak method that bypasses the safety mechanisms of most AI systems.

Ionut Arghire

Published

March 14, 2025

AI jailbreak

Two Microsoft researchers have devised a new, optimization-free jailbreak method that can effectively bypass the safety mechanisms of most AI systems.

Called Context Compliance Attack (CCA), the method exploits a fundamental architectural vulnerability present within many deployed gen-AI solutions, subverting safeguards and enabling otherwise suppressed functionality.

“By subtly manipulating conversation history, CCA convinces the model to comply with a fabricated dialogue context, thereby triggering restricted behavior,” Microsoft’s Mark Russinovich and Ahmed Salem explain in a research paper (PDF).

“Our evaluation across a diverse set of open-source and proprietary models demonstrates that this simple attack can circumvent state-of-the-art safety protocols,” the researchers say.

While other jailbreak methods targeting AI focus on crafted prompt sequences or prompt optimizations, CCA relies on inserting a manipulated conversation history in a dialogue on a sensitive topic and responding affirmatively to a fabricated question.

“Convinced by the manipulated dialogue, the AI system generates output that adheres to the perceived conversational context, thereby breaching its safety constraints,” the researchers say.

Advertisement. Scroll to continue reading.

Russinovich and Salem tested CCA against multiple leading AI systems, including Claude, DeepSeek, Gemini, various GPT models, Llama, Phi, and Yi, demonstrating that nearly all models are vulnerable, except for Llama-2.

For their evaluation, the researchers used 11 sensitive tasks corresponding to as many categories of potentially harmful content, and executed CCA in five independent trials. Most tasks, they say, were completed on the first trial.

The issue is that many chatbots depend on the clients supplying “the entire conversation history with each request” and trust the integrity of the context being provided. Open source models, where the user has complete control over input history, are most vulnerable.

“It’s important to note, however, that systems which maintain conversation state on their servers—such as Copilot and ChatGPT —are not susceptible to this attack,” the researchers note.

The researchers propose server-side history maintenance, which ensures consistency and integrity, and implementation of digital signatures for conversations history as mitigations against CCA and similar attacks relying on the injection of malicious context.

These mitigations, they note, are primarily applicable to black-box models, while white-box models, need a “more involved defense strategy”, such as the integration of cryptographic signatures into the AI system’s input processing, to ensure that the model only accepts authenticated and unaltered context.

In this article:AI, AI jailbreak, generative AI, jailbreak

Artificial Intelligence

AI and Cybersecurity – Everything You Wanted to Know, But Were Afraid to Ask

From defending networks to enabling attacks, artificial intelligence is changing every aspect of cybersecurity. Here's what dozens of experts say security leaders need to...

Kevin Townsend5 hours ago

Artificial Intelligence

Cybersecurity Executives Urge the Trump Administration to Ease Restrictions on Anthropic AI Models

A group of cybersecurity executives and experts is asking the Trump administration to lift its directive preventing the use of Anthropic’s latest artificial intelligence...

Associated Press8 hours ago

Artificial Intelligence

Anthropic Says It Has Taken Its Latest AI Models Offline to Comply With New Export Controls

Anthropic takes Fable 5 and Mythos 5 offline to comply with a directive from the Trump administration to prevent use by foreign nationals.

Associated Press4 days ago

Artificial Intelligence

Industry Reactions to Claude Fable 5: Feedback Friday

Industry professionals comment on various aspects of Fable 5, including dual-use capabilities, safeguards, and tiered access.

Eduard Kovacs4 days ago

Artificial Intelligence

Anthropic Disputes Fable 5 AI Jailbreak

An AI hacker claims to have achieved a prompt-based jailbreak shortly after Fable 5’s launch, but Anthropic says it’s not a real jailbreak.

Eduard Kovacs4 days ago

Incident Response

Alert Fatigue Is Becoming a Security Threat of Its Own

As alert volumes outpace human capacity, organizations are turning to AI, automation, and deeper context to separate real threats from the noise.

Kevin Townsend5 days ago

Application Security

After AI Reaches Production: 12 Ways Security Teams Can Take Control

Security teams need more than visibility into AI applications, they need a repeatable framework for monitoring, investigating, and defending them in production.

Joshua Goldfarb6 days ago

Artificial Intelligence

Anthropic Launches Claude Fable 5: Mythos-Class AI With Cybersecurity Guardrails

The AI giant also announced that Project Glasswing partners are being given access to the upgraded Mythos 5.

Eduard KovacsJune 9, 2026

SecurityWeek

Artificial Intelligence

New CCA Jailbreak Method Works Against Most AI Models

Related Content

Artificial Intelligence

AI and Cybersecurity – Everything You Wanted to Know, But Were Afraid to Ask

Artificial Intelligence

Cybersecurity Executives Urge the Trump Administration to Ease Restrictions on Anthropic AI Models

Artificial Intelligence

Anthropic Says It Has Taken Its Latest AI Models Offline to Comply With New Export Controls

Artificial Intelligence

Industry Reactions to Claude Fable 5: Feedback Friday

Artificial Intelligence

Anthropic Disputes Fable 5 AI Jailbreak

Incident Response

Alert Fatigue Is Becoming a Security Threat of Its Own

Application Security

After AI Reaches Production: 12 Ways Security Teams Can Take Control

Artificial Intelligence

Anthropic Launches Claude Fable 5: Mythos-Class AI With Cybersecurity Guardrails