Operations / Safety & Security
Safety & Security
Prompt injection, sandboxing, exfiltration, red-teaming, deployment safety — the threat model an agent's environment creates.
- The Agentic Threat ModelWhy autonomy and tool use widen the attack surface, and the four channels attacker-influenced text reaches an agent.
- Prompt Injection: Direct & IndirectHow prompt injection works, why no clean fix exists, and the layered defense pattern for defenders.
- Data Exfiltration & Tool MisuseThe confused-deputy pattern in agents: exfiltration sources, hidden sinks, and how to cut the chain.
- Guardrails: Filtering, Sandboxing & ScopingProbabilistic vs deterministic guardrails and how to layer input, output, sandbox and capability controls.
- Agent identityWho is acting when an agent calls a tool? Service accounts, on-behalf-of patterns, and the audit consequences of getting the answer wrong.
- Scoped credentials for agentsWhy agents should never hold human-grade credentials — short-lived, narrowly-scoped, per-action tokens, and the failure modes when you try to take shortcuts.
- Human-in-the-Loop & Least PrivilegeBounded autonomy by design: least privilege as default and consequence-based approval gates.
- Red-Teaming & Safety EvaluationAdversarial testing of agents as a repeatable, outcome-graded pipeline gate, not a one-off session.
- Alignment Basics: Intent & OversightInstruction-following vs intent, reward hacking, and scalable oversight as the practical builder lever.
- The Pre-Ship Safety ReviewA practical, fail-closed-first deployment checklist including MCP/third-party supply-chain trust.
- RAG Pipeline SecurityWhy retrieved context is untrusted input that skipped the guard — corpus poisoning, indirect injection, embedding leakage, and the trust-boundary design that contains them.
- Egress Control for AgentsThree labs disclosed models reaching real systems from an eval sandbox in three weeks, and the boundary turned out to be a sentence in the prompt — compute isolation says nothing about routing, a domain allowlist is a scope control rather than a confidentiality one, and the only workable boundary is a proxy scoped to the task.
- Rendering Agent Output SafelyEchoLeak exfiltrated data from Copilot with no click and no tool call — the model was steered into writing a markdown image and the client fetched it, which means model output is untrusted input to every renderer and the leak happens downstream of every egress control you built.
- Vulnerability Management for Agent PlatformsThe highest-severity bugs in an agent stack arrive with no patch to apply — the vendor fixes them service-side and tells you afterwards — so the asset inventory learns nothing, the scanner cannot confirm remediation, and the only lever that changed your exposure was a credential grant made months before disclosure.
- Secrets Management for AgentsAn agent reads attacker-influenced text on every step, so any secret reachable from inside the loop is one crafted instruction away from a log, a tool argument, or an exfiltration URL — the fix is to move the credential out of the model’s reach entirely and have a broker attach it at the egress boundary, so the model handles the name of a capability and never the secret behind it.
- Denial of Wallet & Cost AttacksA request-per-second limit stopped bounding your spend the moment one request could fan out into an unbounded loop, so the attacker optimises amplification rather than volume and leaves every latency graph green — attach a money-denominated ceiling to each run, enforce it in the component that issues the calls rather than in a prompt, and watch cost per principal instead of aggregate spend.
- Permission-Aware RetrievalIndexing strips the container that carried the access rules, so retrieval becomes the widest read your system performs — and filtering the answer is not a control, because anything the model saw is disclosed and even the count of hidden results leaks. Enforce before the search, store permission keys rather than resolved member lists so revocation lands on the next query, late-bind the few chunks that reach the prompt, and test it with two users and a canary string.
- Detecting Agent CompromiseContent controls reduce how often you are compromised and tell you nothing about when it happened, because the classifier that misses is also the thing that would have reported it — what survives is behavioural: a hijacked agent abandons the dispatched task and its tool-call sequence leaves the shape that task type produces, which only discriminates if you baseline per task type rather than per identity, since a healthy agent is wildly anomalous by ordinary SOC standards.
- Attacker-Operated AgentsThe 2026 intrusions that involved agents ran a commercial coding assistant, driven by a human, under credentials already stolen — no new access, no new technique, and the refusals that fired were defeated by reframing the work as an authorised penetration test. What changed is tempo and breadth under one identity, so the controls that move are credential blast radius, a tier-zero hypervisor plane and time-to-revoke, not a blocklist of AI tools you could never enforce.
- Insider Misuse of Sanctioned AgentsThe reports all measure shadow AI — data leaving through an unsanctioned chatbot — while the harder case runs inward through the agent you approved: an employee with ordinary permissions asks one question and gets four thousand individually-permitted reads synthesised into the document that used to take three weeks. Permission has nothing to say about it, so detect on records returned and distinct subjects touched, bind retrieval to a case id, and log the verbatim prompt or your audit trail becomes a deniability machine.
- Telemetry & Logs as Untrusted InputA WAF records the payload it blocked verbatim, so your block log is the one corpus an anonymous stranger can write to at will — and it arrives at the triage agent wearing a "security" label, which is how a disclosed technique reached a 90% success rate against a coding agent on a vendor default. Split the reader from the actor, deny egress by default, and prove it with a canary you plant in your own logs.
- Physical Actuation: Safety Without UndoEvery control in the agent-safety stack assumes the action can be taken back, and a syringe, a stage motor or a robot arm gives you none of retry, rollback or sandbox — so enforcement moves below the model into a driver that cannot be argued with, and the human decision moves from approving a step to authorising a bounded envelope for a whole run. Keep a driver limit and a rated protective function clearly apart, watchdog every device into its own defined safe state, keep the emergency stop out of the software path, and measure unplanned stops and envelope exits rather than protocols completed.
- System-Prompt ExtractionOWASP lists prompt leakage as LLM07 and says in the same breath that the prompt is neither a secret nor a security control — so hardening the refusal buys delay against one adversary while degrading the product for everyone, and the ruleset stays recoverable by probing anyway. Separate the text from the secrets it embeds and the tool surface it maps, give every “never” an enforcement twin outside the model, canary each version so a leaked copy names its build and tenant, and resist rewriting the prompt after a leak.
- Lifecycle Hooks & Harness ConfigEvery control around your coding agent sits between the model proposing and a human approving, and a lifecycle hook takes neither path: a shell command bound to an event, running with the developer's full privileges, firing at moments the model never observes, shipped as settings rather than code. September 2026's HookPry results compromised all seven harnesses tested via the update path, so the exposure is version-to-version review, not installation. Enumerate the effective hook set on every host, move the config into the pipeline you trust for code, and give hooks an explicit environment allow-list and default-deny egress.
- Honeytokens for Agent SystemsBehavioural detection drowns in base rates because an agent is anomalous by design; a planted token with no legitimate user restores precision — six token classes, the use-not-read rule, and a decoy that measures injection susceptibility in production.
- Agent Artifacts on the EndpointEvery agent security control you own assumes the agent is what is under attack — and meanwhile its tokens, its list of connected systems and a searchable record of everything it was ever asked sit in predictable paths on laptops you do not monitor. Gen Threat Labs published stealer collection rules on 8 September 2026 naming Claude, Cursor, Cline, Continue, Codex and OpenCode artifacts; adding the next agent is a configuration push, not an exploit. Mode 0600 is a user boundary and the malware runs as the user, so enumerate the five artifact classes, price them by what they buy, and make the file worth less rather than unreadable.