Rule-library size decides nothing here, and neither does detection versus prevention. The axis that determines whether a runtime sensor is useful in front of an agent is whether it can attribute a syscall to a tool call — because the entire Kubernetes runtime-security model assumes one workload has one behavioural baseline, and a coding agent's baseline is "anything a developer might do." Get that wrong and you buy a sensor that alerts on every curl the agent was supposed to run. Get it right and the second decision is the one nobody frames properly: killing a tool subprocess does not stop an agent, it hands the loop an unexplained crash and a reason to retry differently.
At a glance
Four projects, four different answers to "and then what".
| Project | Origin & status | Mechanism | Strongest at |
|---|---|---|---|
| Falco | Sysdig, 2016; donated to CNCF 2018; graduated 2024 | eBPF driver (or kernel module) feeding a YAML rule engine over tracepoints and kprobes | Breadth. The largest community rule library, ATT&CK-mapped, and the detections everyone else is compared against |
| Tetragon | Isovalent / Cilium, CNCF | eBPF-native sensors plus TracingPolicy; an override action that stops a syscall before its body runs |
Process lineage and in-kernel enforcement, at the lowest measured overhead of the four |
| Tracee | Aqua Security, open source | eBPF-native sensor with a built-in signature library plus Rego policy, and artefact capture | Forensics. Container escape, injection and privilege-escalation signatures, with the evidence attached |
| KubeArmor | AccuKnox, 2021; CNCF Sandbox, applying for Incubation | AppArmor, SELinux or BPF-LSM policy so the kernel denies the operation; eBPF for visibility | Declarative least-permission. An allow-list of binaries, paths and destinations, enforced inline |
That spread matters more for agents than for ordinary services, and for a mechanical reason: these sensors cost per event, and an agent harness execs a new process for every tool call. A service that starts once and serves requests generates almost no exec events. An agent running a twenty-step task with a shell tool generates hundreds, plus every child process of every command it runs. Whatever the absolute numbers are on your kernel, your agent nodes are the ones where a per-event sensor gets expensive.
Why the standard model breaks on an agent workload
Kubernetes runtime security is built on an assumption that is almost always true and is false here. The assumption is that a workload has a describable baseline: this image runs these binaries, opens these paths, connects to these services, and anything else is worth an alert. That is what makes a shared rule library possible — "a shell was spawned in a container" is a good detection because in a normal container it is a good detection.
An agent inverts it. A coding agent's job is to run a shell, install packages, clone repositories, execute test suites, spawn compilers and open network connections to registries. Point Falco's default ruleset at a coding-agent node and you will trip Terminal shell in container, Launch package management process, Write below etc and a dozen others on the first successful task. The usual response — suppress those rules for that namespace — leaves you with a sensor watching an agent with most of its detections disabled, which is worse than not deploying it, because now there is a dashboard implying coverage.
So the question to ask a candidate sensor is not "what does it detect" but "what can it tell apart". Three capabilities carry the weight:
- Process lineage. Can a rule express "this
curlis a descendant of the agent's shell tool at depth three", as opposed to "acurlhappened in this pod"? The agent's own tool invocations and a payload's subprocess look identical at the pod level and completely different in the exec tree. - Per-process policy, not per-image policy. An agent pod usually runs one long-lived harness process and a stream of short-lived children. A policy scoped to the image covers both with one rule; a policy scoped to a process subtree can allow the harness to do things and constrain the children.
- A run identifier that survives into the event. None of the four gives you this natively, and it is the gap worth engineering around: your harness knows which run and which tool call is executing, and the sensor knows which PID did what. Joining them means stamping something the kernel can see — a cgroup per run, a dedicated uid or gid per session, or a process-group label the sensor exports — and it turns a stream of anonymous syscalls into attributable actions. This is the same join argument as in isolated runs that find each other, applied one layer down.
Practical consequence for the shortlist: Tetragon's strength is exactly the first two, because process lineage is its native model rather than an enrichment. That is a better reason to pick it for agent workloads than its overhead number, which is the one everybody quotes.
Where each one actually attaches
Falco
Falco is the reference point and deserves to be. Created at Sysdig in 2016, donated to the CNCF in 2018 and graduated in 2024, it reads a stream of kernel events through an eBPF driver and evaluates them against rules in a readable YAML syntax, with the broadest community library available and ATT&CK mappings maintained. Its model is alerting: a match produces an event, and doing something about it is a separate component in the ecosystem rather than part of the engine. For agents, treat it as your discovery tool — run it in alert-only mode over a fleet of agent nodes for a fortnight and the output is a behavioural inventory of what your agents actually do, which is the thing you needed before you could write a policy for anything else on this list.
Tetragon
Tetragon is eBPF-native and built around the process tree, which is the property that matters here. A TracingPolicy can match on kernel functions and arguments, and its override action stops the call before the syscall body executes — genuinely preventive rather than reactive. Its own documentation is careful about the distinction that most coverage flattens: the override prevents, while a Sigkill action terminates the process but does not guarantee the in-flight operation was stopped. Hold onto that sentence; it is the whole of the next section. It also measured lowest of the three sensors in the comparison above by a wide margin, which for high-exec-churn agent nodes is not a rounding detail.
Tracee
Tracee is the forensics answer. It ships a signature library with ATT&CK mappings, supports Rego policy for custom detections, and its differentiator is artefact capture — when something fires, you get the evidence rather than a line in a log. It is notably strong on container escape and process-injection patterns, which is the right shortlist for an agent that is running untrusted code by design. If your failure mode is "an incident happened and we reconstructed almost nothing", which is the position described in only one side could see the breach, this is the project that addresses it directly.
KubeArmor
KubeArmor is the odd one out and the best fit for a specific job. It does not use eBPF for enforcement at all — it compiles policy down to AppArmor, SELinux or BPF-LSM so the kernel itself denies the operation, using eBPF only for visibility. Created by AccuKnox in 2021, it has been at CNCF Sandbox since November 2021 and is applying for Incubation. The shape it wants is a declarative allow-list: these binaries may execute, these paths may be written, these destinations may be reached. For an agent whose legitimate tool surface you can enumerate — a support agent, a data-analysis agent, an ops agent with six scripts — this is very close to the right control, and it is the same posture OpenShell takes with Landlock and seccomp, differing mainly in who owns the policy.
The response action is the decision, and "kill it" is usually wrong
For a conventional workload, detection and prevention sit on a straightforward axis: alerting is cheaper and safer, blocking is stronger and riskier, and killing a process is the blunt version of blocking. For an agent, that axis bends, because the thing you are responding to is not the process — it is a loop that will observe the outcome and act again.
Kill the subprocess and the agent sees a tool that exited without output. It has no idea why, it has no reason to believe the action was forbidden rather than flaky, and its trained response to a flaky tool is to try again, often by a different route. You have suppressed one attempt and taught the loop nothing, at the cost of a signal that looked like enforcement on your dashboard. Worse, per Tetragon's own caveat, the operation already in flight when the signal landed may well have completed.
Deny the syscall and the shape changes entirely. The tool gets EPERM, your harness surfaces that as a tool error, and the model reads a sentence describing a refusal. Now it can adapt, ask, or stop, and your trace contains an attributable refusal tied to a named rule rather than an unexplained exit code. That is the difference between a control the agent can reason about and a control it can only stumble into, and it is the same argument as structured refusal and why-trails made at the kernel boundary instead of the tool boundary.
- Prefer deny-before-execute for anything you can name. LSM-hook enforcement (KubeArmor, or Tetragon's override) returns an error instead of a corpse.
- Reserve the kill for the agent process, not the tool. If the conclusion is "this run must stop", stop the run — the harness, the loop, the whole thing — and record why. Killing a child is a half-measure that reads as a full one.
- Make sure the denial reaches the model as text. A policy denial that surfaces as an empty stderr is functionally the same as a kill. This is harness work, and it is the part teams skip.
- Alert-only is a legitimate steady state for agent nodes. Given how wide a legitimate agent's behaviour is, a high-recall detector plus a human queue beats an over-tuned blocker that breaks tasks. Accept that and instrument it properly rather than pretending you will block your way to safety.
When to pick which
| Situation | Start with | Because |
|---|---|---|
| You have agents in production and no runtime visibility at all | Falco, alert-only | Broadest detections and the fastest path to a behavioural inventory; tune later, on data |
| Agent nodes with heavy tool-call churn, and you care about attribution | Tetragon | Process lineage is native, and the per-event cost is the lowest measured of the group |
| You already run Cilium | Tetragon | Same project family, same operational model, one fewer agent on the node |
| The agent's legitimate tool surface is enumerable | KubeArmor | A declarative allow-list enforced at LSM hooks is the cheapest correct control for a bounded surface |
| You need to reconstruct incidents, not just notice them | Tracee | Artefact capture plus escape and injection signatures give you evidence rather than a log line |
| You want one enforcement layer the agent cannot relax | OpenShell or KubeArmor | Both compile to kernel primitives; pick by who should own the policy, the harness or the platform team |
The common production shape is two of these, not one: a high-recall detector in alert mode for the behaviour you cannot enumerate, and an inline allow-list for the part you can. That is not indecision — the two answer different questions, and running only the second one means you never learn what you got wrong.
FAQ
Do I need one of these if I already run OpenShell or a seccomp profile?
They answer different questions. A per-process kernel policy bounds what one agent may do and is installed by whoever launches it. These four are node-wide sensors owned by a platform team, and they see across workloads — including the agent's children, the sidecars and anything that escaped. The overlap is real at the enforcement layer and the visibility is not duplicated.
Will Falco's default rules work on agent nodes?
Not as shipped. Detections like a shell spawning in a container or a package manager running are correct for normal workloads and describe an agent's ordinary Tuesday. Expect to rewrite rather than suppress, and to write rules that reference process lineage rather than image identity.
Is eBPF enforcement safe in production?
Enforcement at LSM hooks is a supported kernel mechanism, and KubeArmor's path through AppArmor and SELinux is decades old. The risk is not the mechanism, it is your policy: a false positive on an inline denial breaks a task. Run any new policy in audit mode first and promote per rule.
How do I get a run or session identifier into the sensor's events?
Stamp something the kernel can see. A cgroup per run, a dedicated uid or gid per session, or a distinct process group are all visible to these sensors and join cleanly to your harness's own trace. Doing this before you write policy is what makes lineage-based rules possible at all.
Which one would have caught an agent exfiltrating through a URL shortener?
None of them, as configured by default — an HTTPS connection to a popular host is not anomalous at the syscall layer. That failure class is an egress-policy problem, not a runtime-sensor problem; the mechanics are in GET-only was a write channel. These tools see the process that made the connection, which is useful afterwards and not a control.
Further reading
On this wiki:
- The Reference Monitor — why the hook point and the trust domain decide whether any of this is a control.
- Sandbox and Isolation Patterns — how node-wide sensors compose with per-process policy.
- Detecting Agent Compromise — what the signal looks like above the syscall layer.
- Structured Refusal & Why-Trails — making a denial readable to the loop that hit it.
- Egress Control for Agents — the failure class these sensors do not cover.
- The tamper-proof half did not ship — the same primitives sold as a platform, and what the DPU layer would add.