Abuse Reports About Your Agent

7 min read

C26
Operation · Governance & Compliance

Abuse reports about your agent.

Somebody you have no contract with is about to email you an IP address, a five-minute window and a sentence about what your software did to their service, and you will have roughly a day to answer before they escalate to your provider, your customer or a blog post. Everything that makes that answerable has to be in place before the message arrives, because it is a reconstruction problem and not a policy problem: unless an outbound request carried an identifier you can reverse-lookup, "was that us, and which run" takes a week and ends in "we think so". Build the reverse lookup first; the policy is the easy half.

STEP 1

Understand what the reporter actually has.

The asymmetry defines the whole exercise. A third party reporting abuse holds a timestamp, a source IP, a user-agent string, a path or an affected resource, and possibly a volume figure. They do not know your tenant model, which product the traffic came from, or that you run agents at all. They cannot name a run, a customer or a prompt.

Your side holds everything about runs and nothing indexed by what they sent. Traces are keyed by run and session; egress is keyed by destination; the NAT or proxy pool that assigned the IP rotates, frequently hourly, and its allocation log is usually the shortest-lived record in the stack. So the first question — is this ours — is a join across three systems owned by three teams, on a key none of them chose.

The report will not use the word "agent". It will say scraping, excessive crawling, unauthorised access attempts, suspicious edits, or a vulnerability-scanning complaint. Classify on the behaviour described, not the vocabulary: an SSRF probe reported as a "security scan" and a crawl reported as a "DDoS" are both routine agent egress seen from outside.

STEP 2

Make outbound requests self-identifying, which is the whole investment.

One afternoon of work decides whether every future report costs an hour or a week. Everything a third party can observe has to carry something you can look up, and it has to be emitted by the HTTP client rather than agreed by convention.

  • A run identifier on the wire. A header on every agent-originated request, opaque to the recipient and resolvable by you to a run, a tenant and a product. This is the single highest-leverage line of code in this page.
  • An honest user-agent with a contact URL. A product token, a version, and a URL that resolves to a page saying what the traffic is and how to reach you. A stranger who can see what hit them will usually write to you before they write to your provider.
  • A request log keyed by what they see. Destination host, source IP, timestamp, run id — queryable by IP and time window, retained longer than your traces and far longer than your IP allocations. This is the index that makes the join possible at all.
  • Egress through a path that logs. A task-scoped proxy is the natural place for all three, and the only place that catches the traffic from code paths nobody remembered.

Where the sanctioned identity mechanism exists, use it: signed-agent schemes let a site verify that traffic is yours cryptographically rather than taking a header's word for it, which converts an attribution argument into a lookup. That is the subject of bot verification and agent access, from the receiving side.

STEP 3

Route the inbox to someone who can answer it.

Most organisations route abuse@ to a security queue tuned for reports about inbound attacks on them. An abuse report is the mirror image — an outbound-behaviour complaint — and the queue that receives it typically cannot query agent telemetry, does not own the egress proxy, and has no mandate over a product team's traffic.

  • Name an owner, and write the runbook before you need it. Receiving, triage, the standing-still default while you investigate, who may speak to the reporter, who can halt a fleet. This is accountability and roles applied to a case with an external clock on it.
  • Publish the route. abuse@, the contact URL in the user-agent, and a page a stranger can find from an IP WHOIS record. Reports that cannot find you do not disappear; they arrive via your provider's trust-and-safety team instead, with you on the wrong side of the conversation.
  • Make a halt reachable in minutes. The default answer during investigation is to stop the behaviour against that origin, which requires the kill switch to be scoped per destination rather than per deployment. If your only lever is "turn the product off", you will not pull it, and the report will escalate.
  • Test the path quarterly. Send yourself a realistic report — an IP, a window, two sentences — and time the answer. The failure is almost never the investigation; it is that the message sat in a queue for three days.
STEP 4

Reconstruct the run, then widen the question.

With the index from STEP 2, the first query is mechanical: IP plus window gives you candidate requests, the run id gives you the trajectory, and the trajectory tells you what the agent was trying to do. Read it before replying, because the interesting finding is almost never the one in the report.

Three things to establish, in order:

  • Was it one run or a population? A single agent behaving unusually is an incident; the same pattern across hundreds of runs is a design default, and the reporter happened to be the first origin large enough to notice. Query by destination across the fleet, not by run.
  • Was the behaviour in scope? Compare what happened against what the run was authorised to do. A crawl that exceeded a budget it never had is a missing control; a probe of a URL the task never mentioned is a task-scope failure and a more serious finding.
  • Did a sanctioned path exist? Bulk APIs, paid data feeds, dumps, a rate-limited partner key. If one did and the agent did not take it, the fix is a client-side route table, because nothing in the loop can know about a contract that is not expressed in code.

Agent abuse is a cumulative property, which is why per-request monitoring never surfaces it: every individual request was well-formed, authorised and successful. The aggregation that would have caught it — requests per origin per hour, summed across the whole fleet — belongs to neither the agent team nor the network team, and is consequently computed by neither. Build that one dashboard and most future reports arrive from your own alerting first.

STEP 5

Reply in a way that closes the case.

Reporters escalate in proportion to how unanswered they feel, so the reply is a control and not a courtesy. Four things, and no more: acknowledge within hours with a case reference, confirm or deny attribution plainly once you know, state what you have stopped and when, and say what changed so it cannot recur.

  • Do not send a determination you cannot support. "We have no evidence this traffic originated with us" is defensible when your index is thin; "this was not us" is the sentence that ends badly when the reporter has a header you forgot you emit.
  • Separate what you fixed from what you are considering. A per-origin budget shipped today is worth more to a reporter than a roadmap item, and it is also the thing that stops the second report.
  • Keep the case record. Report, evidence, timeline, decision, and the change that followed — in the same place as your audit trail. A pattern of reports is a fact about your deployment, and the second regulator or customer question will be about the pattern, not the case.
  • Check the regulatory branch, briefly. Most abuse reports are not notifiable events. Some — unauthorised access to a third-party system, personal data pulled where it should not have been — cross into serious incident reporting or breach-notification duties on a clock you do not control. Make that assessment explicitly rather than by omission.
STEP 6

Treat the report as the external detector you did not build.

A third party found a behaviour your telemetry rated a success. That is the finding, independent of the specific complaint, and it should change two things rather than one.

  • Add the detector. Whatever signal the reporter used — error rate from one origin, request volume, edit frequency, an unusual path — is computable from your own egress log. Every report should leave behind an alert that would have fired first.
  • Make the fix structural, not behavioural. A prompt instructing the agent to be polite is not a control; a per-origin request budget held by the client, which the agent cannot raise, is. The same applies to identification, route selection and backoff: these are properties of a library, not of a model's intentions.
  • Feed it back as a regression test. Add the pattern to your failure taxonomy and a case to the eval set, so the next model or harness change cannot quietly reintroduce it.

If you do exactly one thing from this page: stamp a reverse-lookupable run identifier on every outbound HTTP request your agents make, store it in a log indexed by destination and time, and keep that log longer than you keep traces. It costs an afternoon. It is the difference between answering a stranger in an hour with a run id and a timeline, and auditing your entire fleet in the dark while a deadline you did not set runs down.