AI Blog

The first agentic breach arrived as paperwork — and the form has no field for it

Every public sign that agents are being used to attack people has come from the attacker's side of the wire. Spain's AEPD broke that pattern with a breach notification filed by the victim — compelled, defender-side, adversary-independent evidence, which is the only kind that could ever produce a base rate. The agency's own caveat is the story: one notification is not a trend, and the register it landed in has no field that would make a thousand of them one either.

By Agentic AI Wiki 13 min read

Every public sign that agents are being used to attack people has come from the attacker's side of the wire — a lab's disruption report, a sensor network's telemetry, an affiliate who left a home directory listening on port 8888. Spain's data protection agency broke that pattern on 15 September by publishing something different: a breach notification filed by the victim, under legal compulsion, naming an AI agent as the thing that found the flaw, changed the records and pulled the invoices. That provenance is what makes it worth an article. The agency's own caveat is what makes it worth an argument — one notification does not establish a trend, and the register it landed in has no field that would make a thousand of them one either.

At a glance

Four public evidence sources for agentic attacks have accumulated over the past few months. They are not interchangeable, and the difference that matters is who was obliged to produce them.

Evidence sourceProduced byCompelled?Counted anywhere?
Model-provider disruption reports The vendor whose model was abused No — published voluntarily No shared register
Honeypot and sensor telemetry A security vendor's own network No Per-vendor, per-campaign
An operator's leaked working directory Nobody — an accident No No
A personal-data breach notification The breached organisation Yes — statutory, 72 hours In a register, as prose
Provenance of agentic-attack evidence A matrix of four evidence sources against four properties: compelled by law, produced by the defender, independent of the attacker's cooperation, and structured enough to count. Model-provider disruption reports and sensor telemetry are voluntary and vendor-side; a leaked operator directory is accidental and depends entirely on attacker carelessness; only a statutory breach notification is compelled, defender-side and adversary-independent, and it is the one source with no structured field to count. Provenance of agentic-attack evidence Compelled Defender-side Adversary-independent Countable Model-provider disruption report Voluntary Vendor Own API only No Honeypot / sensor telemetry Voluntary Third party Yes Per vendor Leaked operator directory Accidental Neither Needs sloppiness No Statutory breach notification 72 hours The victim Yes Prose only What a base rate needs Required Required Required Required Holds Partial Fails
Only one row is compelled, defender-side and adversary-independent at once — and it is the row with no structured field for the thing we want to count.

What the notification actually says

The Agencia Española de Protección de Datos describes a single intrusion carried out through an AI agent driven by what it calls a well-known large language model. The agent began by probing generic files for weaknesses, achieved a valid login, and then — inside the authenticated session — searched the application for vulnerabilities on its own. Having found one, it modified personal data and reached invoices.

What the agency withholds is as informative as what it publishes. The affected organisation is not named. The model is not named. The AEPD notes that the account comes from the notifying organisation and still has to be analysed, which is the correct posture for a regulator holding one party's version of events written under a 72-hour deadline.

Read as tradecraft, there is nothing here that a competent human intruder could not have done, and the wiki has made that argument before about an affiliate renting a coding agent: the agent removed skill and hours from the middle of an attack rather than widening what the attack could reach. Two details are still worth pausing on. The reconnaissance happened after the login, not before it, which means it ran inside whatever the credentials were entitled to see and produced traffic that looked like an unusually thorough user. And the outcome was a write — personal data modified — which is very likely why anyone noticed. A session that only reads leaves a much quieter trace, a point the wiki laboured recently over a home-automation server whose four read tools carried the risk.

The intrusion as the victim's logs recorded it One authenticated session containing four phases in sequence: probing generic files, a valid login, an autonomous vulnerability search inside the application, and then modification of personal data plus access to invoices. A detection threshold line sits at the bottom, crossed only by the write; the reads and the reconnaissance sit below it, inside the entitlements the credentials already carried. The intrusion as the victim's logs recorded it Probe generic files Pre-auth, noisy One authenticated session — one identity, one set of entitlements Valid login Nothing to detect Autonomous search for vulnerabilities Reads only — looks like a thorough user Modify personal data The write — this is what surfaced Access invoices A read — quiet Detection threshold actually configured: writes only
Four phases, one session, one identity. Only the write crossed a threshold anybody had set.

Why a compelled filing is different evidence

It samples victims, not vendor footprints

A model provider's abuse report is a census of the abuse that happened to run through that provider's API and survive that provider's detection. It is genuinely valuable and it is structurally partial: it cannot see a self-hosted open-weights model, a competitor's API, or an attack that the provider's classifiers missed. Sensor telemetry has the mirror-image bias — it sees what touched the sensor. A breach register samples the other population entirely. Whoever was breached files, whatever model was used, whether or not a vendor ever learned of it.

It does not require the adversary to cooperate

Most agentic-attack evidence to date has depended on the attacker being sloppy, chatty, or routed through a provider that kept logs. The Aurora case reached the public because an operator left a directory world-readable. That is a sampling method with a strong selection effect: it counts the careless. Article 33 counts the caught, which is a different bias but a more stable one, and one that regulators already spend real money maintaining.

It is the only stream with a legal duty attached

Voluntary disclosure is a business decision, and the incentives run in one direction. A statutory notification duty with a 72-hour clock is not. That is why this filing is the seed of the field's first honest denominator — and why the next section matters more than the incident does.

The counting problem

The AEPD's caution is the most quoted line in the coverage and the least examined: a first notification does not allow a statistical trend to be established. That is obviously true of one data point. The structural version is worse, and nobody has said it out loud.

A breach notification is a form. Its fields are the nature of the breach, categories and approximate number of data subjects, categories and approximate number of records, likely consequences, the measures taken, and the DPO's contact details. There is a free-text narrative. There is no field that says an autonomous agent did this. The agent-ness of this incident reached the public because a regulator chose to write a blog post about one filing, not because a checkbox aggregated it.

Follow that through and the implications are concrete:

  • At a thousand filings, the trend is still not establishable — not for want of volume, but because the distinguishing attribute lives in prose that some notifiers will include and most will not. You cannot query a narrative field for a category nobody was asked about.
  • The first counts will be wrong in both directions. Under-count, because an organisation that cannot tell an agent from a fast human will describe an ordinary intrusion. Over-count, from the moment "an AI agent did it" becomes the reflexive attribution for anything that moved quickly — and it will, because it is a better-sounding sentence to put in front of a regulator than "we had an unpatched authorisation bug".
  • Attribution here is inference, not forensics. The notifier did not read the attacker's model card. They inferred an agent from tempo and session shape. That inference can be right and still be unfalsifiable from the victim's logs alone, which is exactly the evidentiary problem the wiki described in attacker-operated agents: what changes observably is tempo and breadth under one identity, and tempo is not a signature.

The honest read: this filing is a signal about the world and a warning about our instruments. It tells you agentic intrusion has reached real processing of real personal data. It tells you nothing yet about how often — and the machinery that would answer "how often" does not currently exist in any of the regimes that collect the reports.

Three clocks, three recipients, no shared flag

Three reporting regimes, three clocks, no shared flag Three columns comparing the personal-data breach regime, the AI Act serious-incident regime and the network-security regime. Each has a different deadline, a different recipient and a different form, and all three record the cause in free text with no enumerated field for whether an autonomous agent executed the incident. One intrusion, three filings Personal-data breach AI Act serious incident Network-security incident 72 hours from awareness Outer bound 15 days, less for the worst Early warning in 24 hours To the data protection authority To the market surveillance authority To the national CSIRT Its own form Its own form Its own form Cause recorded as free text — no enumerated field for "an autonomous agent executed this"
The same intrusion can start three clocks. None of the three forms has anywhere to write down that an agent ran it.

An organisation in scope for more than one regime files more than once, to different authorities, on different deadlines, and the reports do not join. A personal-data breach runs on 72 hours to the data protection authority. A serious incident involving a high-risk AI system runs on a much shorter clock for the worst outcomes and an outer bound of fifteen days. A significant incident under the network-security regime wants an early warning within 24 hours. The wiki's serious-incident reporting entry works the deadline arithmetic properly; the point here is narrower and more annoying.

Each of those regimes is building its own corpus of incidents. None of them has a common identifier for the incident, a common taxonomy for the cause, or a field for autonomous execution. So the first real dataset on agentic attacks is being written right now, in three incompatible formats, in prose, by people under a deadline who are mostly trying to avoid saying anything that will be used against them later. That is not a reason to be cynical about the regimes. It is a reason to notice that adding one enumerated field — was an autonomous agent involved in executing this incident: yes / no / unknown — is the cheapest possible intervention anyone in this field could make, and it has to happen before the volume arrives, not after.

What to do with this, if you are the one filing

Write the agent-ness into the narrative, with observations rather than conclusions

Nobody is going to ask you. Put it in anyway, and put in what you saw rather than what you concluded: requests per minute inside the session, the ratio of distinct endpoints touched to endpoints your median user touches, whether the reconnaissance preceded or followed authentication, whether parameter fuzzing showed adaptation between attempts. "We assess an automated agent executed this, on the following observations" survives scrutiny. "An AI agent attacked us" does not, and it is the sentence that will poison the base rate for everyone.

Make sure you can still answer the form's actual questions

The compelled fields are the hard ones for an agent-driven intrusion, because they ask what was touched. Categories and approximate number of data subjects means you need per-record read provenance for an authenticated session that read broadly and wrote narrowly. Most teams have writes and have no idea what was read. That gap is an audit-trail problem you cannot fix inside 72 hours, and it is the single most expensive thing about this incident class.

Shorten the session, not the perimeter

The reconnaissance here happened after a successful login. Perimeter controls had already passed. What limits this attack is how much one authenticated session is entitled to reach and how quickly an anomalous one is cut — which is a blast-radius question, not a detection one. If your answer to the agentic-attack briefing is a new detection product, you have bought a signature for something whose only distinctive property is speed.

If you do one thing this quarter: take your last three penetration-test reports and ask, for each finding, whether the exploitation path would have been discoverable from inside a low-privilege authenticated session by an attacker with unlimited patience and no hourly rate. The findings that fail that test are the ones whose cost structure just changed — and the answer is per-session authorisation and rate-shaping, not another sensor.

FAQ

What did Spain's AEPD actually report?

The first notification it has received of a personal-data breach executed through an AI agent. The agent probed generic files, logged in successfully, autonomously searched the application for vulnerabilities from inside the session, then modified personal data and accessed invoices. The affected organisation and the language model used were not disclosed.

Does this mean AI agents are now a major source of breaches?

No, and the agency says so directly: one notification does not establish a statistical trend. What it establishes is that agent-executed attacks have reached real processing of real personal data, which is different from a lab demonstration or a vendor's threat report.

Why does it matter that the report came from the victim rather than a vendor?

Because it is compelled, defender-side and independent of the attacker's cooperation. Vendor reports see only what crossed that vendor's own infrastructure; leaked operator directories count the careless. A statutory breach register samples whoever was breached, regardless of which model was used or whether any provider ever found out.

Could the organisation be wrong about it being an agent?

Yes. Attribution here is an inference from tempo and session shape, not forensic proof. That is a reason to record observations rather than conclusions in the filing, and a reason to expect the early counts to be wrong in both directions.

What would make agentic incidents countable?

One enumerated field in the notification forms — whether an autonomous agent was involved in executing the incident, with an explicit "unknown" option — added before the volume arrives. Everything else is a free-text search over prose that most notifiers had no reason to write.

What is the practical control for this attack shape?

Limiting what a single authenticated session can reach and how fast it can reach it. The intrusion's distinctive property was that vulnerability discovery happened cheaply inside a valid session, so the leverage is in per-session authorisation, rate-shaping and anomaly cut-offs rather than in perimeter detection.

Further reading

On this wiki:

Sources: