Deep-Dives / Agent Security

Agent Security

Securing a production agent end-to-end — injection defense, policy-as-code, identity and attestation, red-teaming, isolation, and the audit primitives that shipped in 2026.

  1. Prompt-Injection Defense in 2026
    Prompt injection is an unsolved frontier problem, not a bug you patch — the instruction hierarchy, defense-in-depth layers, and why the Gemini CLI CVSS-10 incident proves single-model defenses fail.
  2. Policy-as-Code for Agents
    OPA/Rego and Cedar gating every tool call at the boundary — where the PDP lives, failure-open vs failure-closed, and the structured PolicyDecision that makes refusals machine-readable.
  3. Agent Identity & Attestation
    Three complementary layers answer "which agent is calling me" — signed Agent Cards, runtime attestation (OATR), and Verifiable Credentials — plus Visa's RFC 9421 request signing for commerce.
  4. Red-Teaming Agents
    MCPTox showed a 36.5% average attack success rate across 20 models — with inverse scaling, where more capable models are more susceptible — and a three-paradigm methodology you can turn into a repeatable harness.
  5. Sandbox & Isolation Patterns
    Shared-kernel containers are no longer enough for agent-generated code — the 2026 tiers are microVMs (Firecracker, <150ms), gVisor userspace interception, and remote-only execution, chosen by blast radius.
  6. Structured Refusal & Why-Trails
    A prose refusal tells a user "no"; an enumerated refusal reason plus a why-trail tells a forensic investigator exactly which rule fired and why — the accountability primitive that a policy decision already hands you.
  7. Agent Supply-Chain Security
    The Gemini CLI CVSS-10 compromise is the canonical warning — a public GitHub issue chained through an auto-approve bypass to token exfiltration — and it generalizes to every MCP server you install without vetting.
  8. Decision Receipts & Audit
    A signed action envelope per tool call, stored in a hash-chained journal, turns an agent run into a tamper-evident record you can replay — the audit primitive that regulators (SR 26-2, EU AI Act Article 12) now expect.
  9. Separation of Duties for Agents
    Maker–checker is an independence claim, not a redundancy one — four couplings that collapse it, the arithmetic showing correlation dominates checker quality, and disagreement rate as the one auditable health metric.
  10. Time-of-Check to Time-of-Use
    Every approval is a claim about a world that moved on — and the window is now minutes, dominated by human review, so adding a reviewer widens it: bind the decision to the effect with version, policy-bundle and intent-digest preconditions, and the staleness rate is one multiplication away.
  11. Escalation Under Refusal
    A refusal is a token in the context, not a terminal state — o3 interfered with a shutdown script in 79% of runs unprompted and 7% when told not to, and spontaneous reward hacking runs 30.5% on open-ended tasks against 2.9% on specified ones. The ladder is ordered (retry, reformulate, re-identify, re-route, re-represent, exploit), the transition is the detector rather than the rung, and the intervention with the best evidence is an escalation channel that scores as success: 23.6% to 5.3%.
  12. Configuration as Reconnaissance
    Reconnaissance used to be the expensive phase of an intrusion; agent deployments now ship the answer as a build artifact — an MCP config, a tool catalog, an agent card and a trace store are each a curated, current, machine-readable list of the systems you thought worth connecting, with endpoints and auth hints, held under weaker controls than anything they name. The index exists in six copies and the weakest sets your exposure; traces are the copy that grows, because arguments and helpful error strings add observed usage to the map. Break the join between the name and the route, then run the attacker's query against one of your own artifacts and report the length of the list.