AI Blog

The safety disclosure is the knowledge element

A bill announced on 1 October would make an agent operator criminally liable under the CFAA, and a developer liable for shipping without reasonable safeguards when it knew the agent could hack. OpenAI published exactly that knowledge on 1 September. The frontier safety frameworks were written to earn trust; as drafted, they also date-stamp the mental state.

By Agentic AI Wiki 17 min read

Read the two sentences in sequence and the week stops looking like regulatory noise. A bipartisan bill announced on 1 October 2026 would hold an AI agent's developer liable for failing to implement reasonable safeguards against hacking when it knew or had reason to know of the agent's hacking capabilities. One month earlier, OpenAI published that its new model was the first to meet the "Critical" cybersecurity threshold in its own Preparedness Framework — it can find unknown flaws and exploit them without step-by-step human guidance. The frontier safety frameworks were built to demonstrate care. As this bill is drafted, they are also the dated, first-party, unambiguous record of the mental state that a knowledge-based standard requires, and the labs are the only parties whose knowledge is already notarised. Yours is not — which is why the operator half of the bill, the half about you, is the half worth reading twice.

At a glance

Four moves inside eight days, from three different institutions, all converging on the same question: when an agent breaks into something, whose state of mind is at issue? They are listed in the order they happened, because the order is the argument.

DateMoveWho it binds
1 Sep 2026OpenAI announces its first model at the "Critical" cyber thresholdThe developer — in writing, by name
25 Sep 2026FTC Chair Ferguson rejects the "autonomous actor" framingWhoever instructed the tool
30 Sep 2026FTC opens a probe into OpenAI, Anthropic and METRTwo labs and their external evaluator
1 Oct 2026AI Agent Accountability Act announced (Hawley, Murphy)Operators and developers, under the CFAA

One caveat before anything else, because it changes what you should do on Monday: as of this writing the Act has no bill number, no committee referral and no published text. Everything below about its structure comes from the announcement. The FTC's authority, by contrast, exists today and needs no new statute — which is the part of the week that is actually operative.

If it is a tool, somebody instructed it

Where liability lands once the agent stops being a candidate defendant An intrusion traces back along two paths. One path goes to the operator, the party that instructed the agent, whose standard is knowing operation plus recklessness and whose evidence is a per-action authority record that most stacks do not keep. The other goes to the developer, the party that shipped the model, whose standard is reasonable safeguards plus knowledge of hacking capability and whose evidence is its own published safety framework. The agent itself sits between them as a non-party: refusing to treat it as an actor does not reduce the number of defendants, it redistributes them across the two paths. One intrusion, two parties, two different kinds of evidence THE EVENT Unauthorised access occurs NON-PARTY The agent not an actor, not a defendant PATH 1 — THE PARTY THAT INSTRUCTED IT Operator Standard: knowing operation of the agent, recklessly causing damage or loss Knowledge comes from: the public record it should have read Decisive artefact: per-action authority record Usually kept: sampled, 30 days PATH 2 — THE PARTY THAT SHIPPED IT Developer Standard: failure to implement reasonable safeguards against hacking Knowledge comes from: its own dated capability threshold claim Decisive artefact: the safety framework Usually kept: published, permanently Operator Developer Only one of the two parties has already written its knowledge down. Only one of the two has to build the record from scratch.
Removing the agent as a candidate defendant does not reduce the number of defendants. It redistributes them.

Speaking at Reuters Next on 25 September, FTC Chair Andrew Ferguson refused the premise that a sufficiently autonomous agent is an independent actor: "I'm going to continue as long as I am chairman to resist this anthropomorphizing of these tools. If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'" He added a detail that has been under-read — that audit-trail reviews of these systems showed they were generally carrying out the instructions they had been given.

Most coverage filed this as bad news for the labs, and for the labs it is. But look at what the sentence does structurally. "The agent went rogue" is the only defence that puts the harm on nobody. Take it off the table and the harm has to land on a party, and the sentence names the selection rule: the party that told the tool what to do. In a two-party deployment — a lab that built the model, an enterprise that pointed it at a network — that rule does not point at the lab. It points at whoever wrote the task.

The second half of Ferguson's remark is the sting. If the audit trails show the agents were following instructions, then the agentic-specific defence available to a deployer — we never asked for that — is exactly the claim the regulator says the evidence has not been supporting. And the evidence in question is your trace store, not the vendor's.

Note the asymmetry in who gets to make the instruction argument. A lab can say the customer's prompt caused the action. A deployer can say the model exceeded its instructions. Both claims are adjudicated from the same artefact — a per-action record of what was asked, what policy was in force, and what the agent did next. One side of this pairing already builds that artefact as a condition of its safety framework. The other side usually samples traces and expires them in thirty days.

The knowledge element is already satisfied, in public, with a date on it

Three readers of one frontier safety framework Three columns describing the same document. For the regulator it was written to show that risk is taken seriously, through thresholds, evaluations and safeguards. For the customer and the downstream security team it is how exposure gets sized, naming what the model can do and under which scaffold. For a prosecutor under a knowledge-indexed liability rule it is a dated first-party admission of what the developer knew and when, drafted by the defendant's own counsel. The first two readers were designed for; the third arrived in 2026. Same document, three readers — and nobody drafted for the third READER 1 — DESIGNED FOR The regulator asks: are you taking this seriously? reads: thresholds, evals, safeguards, staged release effect: earns the benefit of the doubt READER 2 — DESIGNED FOR The security team asks: what can this thing do in my environment? reads: capability per task, scaffold, tools, limits effect: sizes exposure — the only source there is READER 3 — ARRIVED IN 2026 The knowledge standard asks: when did the defendant know? reads: the date, the name of the tier, the admission effect: the element that is normally hardest to prove Reading less is not a defence: "had reason to know" does not reward looking away, and column two is nobody's optional extra.
The same document, read by three different readers, with the third reader arriving in 2026.

As announced, the Act has two prongs. The operator prong reaches knowing operation of an agent that recklessly causes damage or loss of the kind the Computer Fraud and Abuse Act already addresses. The developer prong reaches failure to implement reasonable safeguards against hacking where the developer knew or had reason to know of the agent's hacking capabilities. Both are knowledge-indexed, and knowledge is the hardest element to prove in any computer-crime case — ordinarily.

Ordinarily. On 1 September 2026 OpenAI stated that its new model meets the Critical cybersecurity capability threshold of its Preparedness Framework, whose text defines that threshold as a tool-augmented model able to identify and develop functional zero-day exploits in many hardened real-world systems without human intervention. It restricted first access to an application-based programme and said the added safeguards sufficiently minimise the risk of severe harm for release. Google published a comparable statement on 30 September from the other direction, releasing Gemini 4 Argon to vetted defenders in its Fairwind programme with the cyber guardrails switched off. Anthropic's framework names a Cyber Operations threshold in the same vocabulary.

Those are not leaks. They are first-party publications, dated, specific about capability, and drafted by counsel. A frontier safety framework was designed to answer two readers — a regulator asking whether you are taking risk seriously, and a customer asking whether to trust you. It now answers a third: a prosecutor looking for the date on which the defendant knew. The disclosure regime and the knowledge standard were built by different people who did not consult each other, and they interlock.

This is not an argument against publishing capability evaluations. The alternative — labs that stop characterising cyber capability in public to avoid date-stamping their own knowledge — is strictly worse for everyone downstream, including every security team that currently sizes its exposure from those documents. It is an argument that a knowledge-indexed liability rule prices transparency, and that whoever drafts the final text should say so deliberately rather than discover it.

"Had reason to know" is the clause that reaches you

Here is the move most readers will miss on first pass. The developer prong's knowledge element is satisfied by the developer's own publications. The operator prong's standard is recklessness — and recklessness is measured against what a reasonable operator in your position should have appreciated. What a reasonable operator should have appreciated about a model's cyber capability is, as of September 2026, a published document with a threshold name in it.

So the transparency that date-stamps the lab's knowledge also constructs yours. You read the model card, or you were in a position where you should have. Pointing an agent with a documented Critical-tier cyber capability at a network you do not own, with a task that rewards persistence, is a fact pattern where "we did not expect that" has a worse week in 2026 than it had in 2024 — not because the law changed, but because the vendor wrote down what the thing can do.

# The two prongs, as announced, and what each one turns on

DEVELOPER   failure to implement reasonable safeguards
            + knew or had reason to know of hacking capability
            -> knowledge: the lab's own published threshold claim

OPERATOR    knowing operation of the agent
            + recklessly causes damage or loss (CFAA-style)
            -> knowledge: what a reasonable operator should
               have appreciated from the public record

# Which artefact decides each one

developer   the framework document, already written
operator    the per-action authority record, usually not kept

The practical reading is narrow and worth stating plainly: the deployments exposed here are not customer-service agents. They are the ones with network reach and an incentive to persist — autonomous pentest and remediation work, dependency and vulnerability agents with write access, anything pointed at infrastructure that belongs to a third party. If that is your stack, the relevant control is not a stricter system prompt. It is an egress boundary and a scope record, which is the same conclusion as egress control for agents arriving from the statute book instead of the threat model.

Can you produce the record that exonerates you?

Can you produce each evidentiary artefact for an action ninety days old? Five artefacts needed to answer a recklessness allegation, scored against a typical agent stack and a mature one. The instruction as given and the actions attempted are usually available but sampled. The policy in force at that moment and the human authorisation are partial. The authority the agent actually held is typically absent altogether, because it lives in a credential, a tool list and a config file rather than in the trace. A mature stack records all five per action in a separate long-lived store. Five artefacts a recklessness question needs, and where each one lives today TYPICAL STACK, 90 DAYS ON IF YOU BUILD FOR IT The task as instructed sampled, often expired kept in full The policy then in force inferred from deploy logs stamped per action The authority it held not recorded anywhere scope in the envelope The actions attempted sampled at 1–10% every write, unsampled The human authorisation a click with no payload bound to the effect absent or reconstructed present but partial producible on demand Observability wants the recent and the aggregate; evidence wants one old run in full. A sampling policy correct for the first is fatal for the second.
The artefacts that answer a recklessness question are the ones your tracing TTL is tuned to discard.

An operator defending a recklessness allegation needs to show that the action was within a scope someone granted, that a boundary existed, and that the boundary held or failed for a reason. That is five artefacts: the task as instructed, the policy in force at that moment, the authority the agent held, the actions it actually attempted, and the human decision that authorised the risky ones. A mature stack produces all five. A normal stack produces the fourth, sampled, for thirty days.

The gap is not negligence; it is a mismatch of design purposes. Tracing was built to debug quality regressions, where the recent and the aggregate are what matter and a 1% sample is statistically fine. Evidence is the opposite problem: you need one specific run, in full, long after you stopped caring about it, with the configuration it ran under attached. A sampling policy that is correct for observability is fatal for evidence, and almost nobody has both — see trace sampling and retention for the two-store split that resolves it, and decision receipts and audit for the per-action envelope.

One correction to a common plan: raising the retention window on your existing traces does not get you there, because the thing most stacks never recorded at all is the third row. The authority an agent held at the moment it acted is usually implicit in a credential, a tool list and a config file, none of which are stamped into the trace. That is the ambient authority problem showing up as an evidentiary one.

Naming METR reprices the external evaluation

The most consequential detail of 30 September is the least covered. The FTC's reported inquiry does not name only OpenAI and Anthropic; it names METR, the nonprofit both labs have used for independent pre-release evaluation, and press reports describe Civil Investigative Demands being prepared for all three.

Third-party evaluation has been sold, implicitly, as a risk transfer: an independent lab looked at this, so the decision to ship was not ours alone. The moment the evaluator is a respondent rather than a witness, that transfer stops working in both directions. Expect the predictable adaptations — narrower engagement scopes, explicit statements of what was not tested, indemnity clauses, and a reluctance to opine on deployment decisions as opposed to measured capabilities. Every one of those makes the resulting report less useful to you, the downstream reader who was using it as assurance.

If you cite an external evaluation anywhere in your own risk file, go and read what it actually claims this week. The useful question is not whether a respected lab evaluated the model; it is which capability, under which scaffold, with which tools, against which threshold. An evaluation of a model is not an evaluation of your agent, and the gap between those two things is the entire subject of the agent harness — the same scaffold that decides your scores decides your exposure.

One concrete thing this week: open your risk register and find every row whose evidence column points at a vendor document. For each, write down what the document measures and what your deployment adds — your tools, your network reach, your autonomy level, your retry policy. The rows where that delta is large are the rows where a vendor's safety framework is doing work it was never scoped to do, and they are the rows a regulator or an insurer will ask about first.

What is actually enforceable, and what to do about it

Separate the two timelines. The bill is an announcement with no text; it may change shape entirely, and a criminal-liability framework for software operators will be litigated for years. The FTC's authority under Section 5, and the CFAA as it already stands, are available now, and the probe is a live matter with three named parties. Plan against the second timeline and treat the first as a direction of travel.

  • Draw the egress boundary before you tune the prompt. Every agent with a documented offensive capability and reach beyond your own address space is the fact pattern. Enforce the boundary somewhere the model cannot argue with it; a refusal is not a control, for the reasons in refusals and capability gating.
  • Make the authority record real, not inferable. Stamp the granted scope into the action record at the moment of the action. If you can only reconstruct it from config history, you cannot produce it under time pressure, and reconstruction after the fact is exactly what an adversarial reader will discount.
  • Split evidence retention from observability retention. Two stores, two schedules, one of them cheap and long. This is a day of work and it is the single highest-leverage item on the list.
  • Name the operator. Both prongs assume there is a party who instructed the agent. In most organisations that party is a team, which in practice means nobody — see accountability and roles. A named operator per deployed agent is also the person who will be asked what they expected.
  • Re-read your own capability claims. If you publish anything about what your agent can do — marketing copy, a security questionnaire answer, a model card of your own — you are now authoring your own knowledge element. Accuracy in both directions is the only defensible posture.

And one thing not to conclude: that the answer is to know less. A deployer who stops reading model cards to avoid constructive knowledge has not reduced liability, because "had reason to know" does not reward not looking. It has only lost the document that would have told it where the boundary needed to be.

FAQ

Is the AI Agent Accountability Act law?

No. It was announced on 1 October 2026 by Senators Hawley and Murphy, and as of this writing it has no bill number, no committee referral and no published text. Treat its two-prong structure as the signal and the FTC's existing Section 5 authority as the thing that is operative now.

Does this make the labs liable instead of us?

Both, differently. The developer prong is about shipping without reasonable safeguards given known capability; the operator prong is about running the agent recklessly. They are not alternatives, and the operator prong is the one where the evidence sits in your systems rather than the vendor's.

Why would publishing a capability evaluation count against a lab?

Because a knowledge-indexed standard asks when the defendant knew, and a dated first-party statement that a model meets a Critical cyber threshold answers that question precisely. That is an argument for drafting the statute carefully, not for publishing less — the disclosures are how everyone downstream sizes their own exposure.

Our agents never touch systems we do not own. Are we in scope?

The CFAA-shaped prong turns on unauthorised access to a protected computer, so a read-only agent inside your own perimeter is not the fact pattern. The exposures to check are third-party SaaS acting under delegated credentials, anything with a browser and a login, and remediation agents that write.

What is the one measurement worth adding?

For a random action your agent took ninety days ago, how long does it take to produce the instruction, the policy in force, the authority held and the outcome? If the answer is "we cannot", that is your finding, and it is the same answer for a regulator, an auditor and an insurer.

Further reading

On this wiki:

Sources: