The Spanish Data Protection Agency (AEPD) has published details of the first notification of a personal data protection breach executed by design through an AI agent.
Investigation into the attack is continuing, and the AEPD uses its words carefully. Nevertheless, although AI-assisted attacks have become common (deepfakes, authoring phishing emails, scaling attacks through automation, etcetera), this appears to be the first known agentic attack outside of a rogue frontier model agent. Bad actor agents are moving beyond a theoretical probability into the real world.
The attack itself involved a successful login, followed by a search for vulnerabilities, and the ability to modify personal data and access invoices. “What is relevant from a data protection perspective,” writes AEPD, “is that a third party would have used an AI agent as an instrument to successfully chain together different phases of the attack.”
This, suggests the agency, is a qualitative change. “An agent can receive a goal, plan intermediate tasks, use tools, execute code, consult sources, interpret results, and modify its actions autonomously, based on what it finds.” And, it should be added, at speed.
The effect requires a four-fold modification to risk management. First, the danger of AI assistance and adversarial agents must become part of risk analysis. Second, incident response times must be improved. Third, the importance of digital IDs and credentials must be recognized, and they must be better protected. And fourth, these modifications cannot be achieved solely through manual intervention.
“Human supervision remains essential, but it must be supported by detection, containment, and response mechanisms capable of operating quickly enough,” says the AEPD – which is a long way of saying that in the adversarial AI era, defense must also be AI assisted, but with a human in the loop.
Commenting on the incident, Simon Phillips, CTO at CyberVerse echoed AEPD’s careful choice of words. “We need to treat this incident with caution and avoid scaremongering the public with stories around AI once again running rogue. We don’t have enough information to understand what happened or how the model carried out this breach,” he said.
“But, the three possibilities that most security experts will consider, include:
- An actor deliberately found a way to bypass the guardrails of a model, potentially through a jailbreak, which enabled them to break into a third party.
- The incident is related to the recent tests carried out by major AI players, including OpenAI and Anthropic, and this is another example of a model escaping a poorly configured testing environment and carrying out autonomous tasks to reach an objective set by a human, but with very little direction from that human.
- A penetration tester has built a model based on a popular LLM, which allowed them to carry out the activity without authorization.”
If the Spanish firm’s notification to its data protection agency is genuine, any one of these scenarios is a possible cause. However, “Out of all these scenarios, the first is the most concerning because it would highlight an actor has been able to bypass the controls enforced by an AI model’s operators,” adds Phillips. “Hopefully we will understand more soon, because organizations need to know what they are facing with AI and where to invest their defenses.”
Is this a blip, a misleading filing with the AEPD, or the expected portent of a more dangerous future?
Related: EU Chief Warns of AI-Powered Hacking, Moves to Rein In Social Media
Related: Hackuity Raises $19 Million for AI-Powered Vulnerability Management
Related: Exein Secures $270M at $1.7B Valuation for Physical AI Security
Related: CISOs Race to Control AI Agents Without Destroying Their Value
