AI Blog

Pydantic AI vs Agno vs smolagents vs Strands: only one of them changes your threat model

Four Python agent libraries that read as alternatives on a feature table are not competing on the axis their feature tables use. Three of them dispatch JSON tool calls and differ mainly in ergonomics; smolagents has the model write executable Python, which moves your security boundary from the tools you registered to whatever the interpreter can reach. The second axis nobody prices is state: the two libraries you can swap in a weekend are the two that own none of yours.

By Agentic AI Wiki 11 min read

Three of these four libraries will produce the same incident report when something goes wrong, because they dispatch the same JSON against the same registered tool table. The fourth has the model write Python and run it, which is not a better or worse choice — it is a different security architecture wearing the same Agent(...) constructor. Pick on that first, and on ergonomics never.

At a glance

All four are permissively licensed Python libraries that will give you a working tool-using agent in about fifteen lines. They diverge on what they consider their job.

ProjectLicenceWhere its weight sitsWhat it declines to own
Pydantic AI MIT The typed contract at every boundary — inputs, tools, outputs. Runtime, persistence, deployment.
smolagents Apache-2.0 The action representation: the model writes Python, not JSON. Nearly everything else — the core is famously about a thousand lines.
Agno Apache-2.0 A runtime and control plane: sessions, memory, traces, approvals. Little. That is the point, and the cost.
Strands Agents Apache-2.0 A model-driven loop with a deployment target already in mind. Opinions about your application layer.
GitHub stars, four Python agent libraries Horizontal bar chart of approximate GitHub star counts in late August 2026: Agno about 41,800, smolagents about 28,900, Pydantic AI about 19,400, and Strands Agents above 2,000. GitHub stars (thousands) 0 12 24 36 48 Agno 41.8k smolagents 28.9k Pydantic AI 19.4k Strands Agents 2k+ Approximate counts, late August 2026. Strands also reports 150k+ monthly PyPI downloads, which is the number that does not appear on this axis at all.
Late-August 2026 counts. Strands is an order of magnitude behind on this axis and reports 150k+ monthly PyPI downloads, which is a reminder that the axis is measuring attention, not use.

Read that chart as four different products rather than one leaderboard. Agno's number includes people who came for a runtime and a UI; Pydantic AI's includes people who already trusted the Pydantic name; Strands' PyPI figure against its star count is the signature of a library that arrives through an employer rather than through a blog post. None of that tells you which one to use.

Where each library places its weight A matrix over four axes — typed contracts, code as action, owning persistent state, and an opinion about deployment. Pydantic AI leans hardest on typed contracts, smolagents on code as action, Agno on owning state, and Strands on deployment. Where each library places its weight Typed contract Code as action Owns your state Deployment opinion Pydantic AI Strong No No None smolagents Light Its whole thesis No Sandbox Agno Medium No Sessions, memory AgentOS Strands Agents Medium No Partial AWS-shaped Where its weight sits Present Deliberately left to you
Each library is strong in exactly one column. The columns are not equally expensive to get wrong.

The dispatch mechanism is the real fork in the road

JSON tool dispatch versus code as action Two paths from a model's output to an effect. In JSON dispatch the framework matches a name against a registered tool table, so the reachable surface is the tools you registered. In code as action the model emits Python that an interpreter executes, so the reachable surface is whatever that process can reach, and the sandbox becomes the boundary. PYDANTIC AI · AGNO · STRANDS SMOLAGENTS CODEAGENT Model output {"name": "search", "args": {…}} Framework dispatch name matched against the registered tool table Reachable surface exactly the tools you registered, and nothing else Model output results = [f(x) for x in load()] Interpreter AST walker locally, or E2B / Docker / Modal / Wasm remotely Reachable surface whatever the process can reach — the sandbox is the boundary
The left path fails closed by construction. The right path fails closed only if you built the sandbox.

Pydantic AI, Agno and Strands all take the ordinary route: the model emits a structured tool call, the framework looks the name up in a table you populated, validates the arguments against a schema, and calls your function. The set of things the agent can do is the set of things you registered. That property is not a feature any of them advertises, because it is inherited from the provider's function-calling API — but it is the reason a review of "what can this agent do" is a review of a list.

smolagents inverts that. Its CodeAgent has the model emit Python, which is then executed — and the framework's own framing is that this is the point, because composition, loops and control flow come free instead of being simulated across multiple round trips. It is a genuinely good idea with real evidence behind it, and it is the choice with consequences.

What "executed" means, concretely

By default that execution happens in LocalPythonExecutor, an AST-walking interpreter rather than a call to exec(): imports are allow-listed and the operation count is capped, which blocks the obvious infinite loop and the obvious import os. The project's own documentation is honest about the ceiling — it is safer than exec(), and it is not a security boundary. For that you attach a remote executor: E2B, Docker, Modal, Blaxel or Wasm.

So the real comparison is not "smolagents vs Pydantic AI". It is "smolagents plus a sandbox vs Pydantic AI", and the sandbox is a running cost, an operational surface and a latency budget. If you are already running untrusted code somewhere, that cost is already paid and smolagents is close to free. If you are not, adopting it means adopting sandboxing as a subsystem, and that decision deserves to be made deliberately rather than inherited from an import statement.

Where each one actually fails

Comparatively, across the same failure: a model that has been talked into doing something it should not. Under JSON dispatch in any of the other three, the damage ceiling is the union of your registered tools — bad, bounded, and auditable from a file. Under code as action with a local executor, the ceiling is whatever the Python process can reach, which on a developer laptop is the developer's credentials and on a server is the service account. Under code as action with a remote sandbox, the ceiling returns to being bounded, and the boundary is now a piece of infrastructure you own and must keep patched. Three postures, one API shape.

The second axis: who holds the state

Set dispatch aside and a different ranking appears, and it is the one that predicts what a migration costs in eighteen months.

Pydantic AI and smolagents own nothing durable. Conversations are yours to store, memory is yours to define, deployment is yours to arrange. That is why they read as "thin", and it is also why replacing either is a weekend: you delete a dependency and keep your data. Pydantic AI in particular is doing one thing on purpose — putting a validated type at every boundary a language model touches, so that a malformed tool argument is a caught exception rather than a mysterious downstream failure. It is the most conservative choice on this page and the easiest to reverse.

Agno is the opposite proposition, deliberately. It is a framework, a runtime and a control plane: AgentOS gives you chat, sessions, traces, evaluations, memory and approvals over your running agents, with memories keyed to a user ID and accumulated across conversations. For a team that would otherwise build all of that, this is enormous leverage and the star count is earned. It is also the one whose adoption reaches your data layer. Once six months of user memories live in its schema and your product surfaces them, you no longer have a library dependency; you have a system of record, and swapping it is a migration with a rollback plan.

Strands sits in between, with its centre of gravity in the deployment story rather than the application one. It presents a model-driven loop and a wide set of provider classes — Bedrock, OpenAI, Anthropic, Gemini, Ollama, Mistral, LiteLLM, SageMaker — which is what a library looks like when it expects to be adopted by an organisation that has already chosen its cloud. The model-agnosticism is real and worth having; the shape of everything around it is still AWS-shaped, and that is a reasonable thing to want when it matches where you deploy.

The inversion worth noticing: the library with the largest star count is the one with the highest exit cost, and the library with the smallest is backed by the vendor most likely to still be maintaining it. Neither fact is an argument by itself. Both are absent from every feature table.

When to pick which

SituationPickBecause
Agent inside an existing Python service, typed end to end Pydantic AI Validated boundaries, no runtime to adopt, trivially reversible.
Data work: the task is calling libraries and composing results smolagents Code is the right action representation — and budget the sandbox as part of the decision.
You need sessions, memory and a review UI, and would otherwise build them Agno Genuine leverage. Plan the data-ownership question before the schema fills up.
Deploying into AWS with Bedrock in the mix Strands Agents Fits the environment; keeps model choice open.
You cannot articulate which column you need Pydantic AI, then re-evaluate It is the choice that costs least to have been wrong about.

One caution that applies to all four: the agent loop itself is not the hard part, and it is not what you are buying. Retry policy, tool dispatch and a message list are a few hundred lines in any of them. What you are buying is a set of decisions about validation, execution and persistence — and the only one of those three that you cannot cheaply revisit later is persistence.

FAQ

Is code as action actually better than JSON tool calls?

For tasks that are mostly composition — filter this, join it to that, loop over the result — yes, and measurably so, because the model expresses in one block what JSON dispatch spends several round trips on. For tasks that are a sequence of discrete, individually-approvable effects, it is worse, because you have replaced a reviewable list of calls with a program. See code as action for the longer argument.

Can I use smolagents without a sandbox?

You can, and for local experimentation against your own data it is reasonable. The AST-walking local executor allow-lists imports and caps operations, which is meaningfully safer than exec(). It is not a boundary you should put untrusted input behind, and the project says so itself. In production, attach E2B, Docker, Modal, Blaxel or Wasm.

Does a bigger star count mean a safer bet?

It means more people arrived. On this page the largest project is also the one whose adoption reaches furthest into your data model, and the smallest is backed by a cloud provider. Stars measure attention; what you want to know is maintenance and exit cost, and neither is on that chart.

How do these relate to LangGraph, CrewAI and the OpenAI Agents SDK?

Those sit further up the orchestration stack, with explicit graphs or handoff primitives for multi-agent work. The four here are lighter: three are libraries you call from your own control flow, and one (Agno) is growing a runtime underneath it. If your problem is coordinating several agents, that comparison is the relevant one.

What if we need to switch later?

Then keep the two things that make switching cheap: your tool functions as plain Python with no framework types in their signatures, and your conversation and memory state in a schema you defined. Do that and any of these four is a few days. Skip it and the framework you picked becomes a fact about your database.

Further reading

On this wiki:

Project sources: