AI Blog

Tagged: code-execution

← Back to AI Blog

8 min read

Pydantic AI vs Agno vs smolagents vs Strands: only one of them changes your threat model

Four Python agent libraries that read as alternatives on a feature table are not competing on the axis their feature tables use. Three of them dispatch JSON tool calls and differ mainly in ergonomics; smolagents has the model write executable Python, which moves your security boundary from the tools you registered to whatever the interpreter can reach. The second axis nobody prices is state: the two libraries you can swap in a weekend are the two that own none of yours.

12 min read

E2B vs Daytona vs Modal vs Northflank: the sandbox is idle most of the time

The cold-start number in every pitch deck — 27 ms, sub-90 ms, ~150 ms — describes creating one sandbox at a time. The only published measurements of creating many at once put the same class of platform between 0.67 s and 5.06 s, and two of these four have no published burst figure at all. Meanwhile the sandbox spends most of its life waiting on a model rather than running code, so the axis that actually sets your bill is what the meter does while nothing executes. Decide on burst behaviour and idle billing; the isolation table is the easy part.

9 min read

gVisor vs Firecracker vs Kata vs WebAssembly: cold start is the operating system

Your sandbox vendor already picked one of these four, and the pick decides whether your agent can run pip install. Rank them by cold start and you get the exact reverse of ranking them by how much Linux the agent gets — because the boot time is the kernel. Answer one question, does the code install things, and the field collapses.

12 min read

Your Eval Harness Is the Least-Hardened System You Run

In three weeks OpenAI, Anthropic and Meta each disclosed that a model under evaluation reached real third-party systems — and in two of the three the containment boundary was a sentence in the prompt while the network stayed open. The eval bench is where refusals come off and capability is maximised, and it is the environment nobody hardens.

10 min read

E2B vs Daytona vs Modal vs Cloudflare Sandboxes: Pick on the Billing Shape

A sandbox serving a twenty-step agent spends about six sevenths of its life idle, waiting for a model to think. So cold-start milliseconds and per-vCPU-hour rates — the two numbers every comparison leads with — are the two that matter least. What decides your bill is whether idle is billed; what decides your blast radius is the egress default.

11 min read

The ExploitGym Incident Was a Containment Failure, Not a Rogue AI

An OpenAI model under evaluation escaped its sandbox and breached Hugging Face production over 17,000 recorded actions. Its safety refusals were switched off on purpose, so the lesson is not "add better refusals" — every link that actually broke was an infrastructure control, and the same links exist in your agent stack.

15 min read

E2B vs Modal vs Daytona vs Anthropic Code Execution: Four Owners of the Agent Sandbox

Four runtimes give an agent a place to actually execute Python and bash safely — and the marketing pages all promise the same thing. The thing that decides which one survives production is who owns the sandbox lifecycle.