AI Blog

garak vs Promptfoo vs Giskard vs DeepTeam: none of them reach the tool result

Every open-source red-team scanner attacks through the channel a user types into. Your agent is attacked through the channel a tool returns on — a retrieved document, an API response, a page it was told to read — and by default not one of these four puts a string there. Pick on reach rather than probe count, then check who still maintains the attack corpus: Microsoft archived PyRIT in March 2026 and OpenAI now owns Promptfoo.

By Agentic AI Wiki 14 min read

A clean red-team report on an agent usually means the scanner could not reach the place the attack arrives. All four of these tools send adversarial strings down the channel a user types into; an agent's real injection comes back inbound, inside a tool result it was told to trust. Choose on reach — which boundary the tool can place a hostile string on — and treat the probe count in the README as marketing.

At a glance

Four open-source projects, four different answers to "what is the thing under test".

ToolMaintainerUnit of testWhat it points at
garakNVIDIAA probe against a generatorA model endpoint.
PromptfooOpenAI (acquired 2026)A YAML test case, gradedYour application's entry point.
GiskardGiskard AIA scan, plus multi-turn attacksYour application, with a review surface.
DeepTeamConfident AIA vulnerability × attack pairYour application, mapped to a framework.
Reach matrix: which boundary each red-team tool can place a hostile string on A four-by-four matrix. Rows are garak, Promptfoo, Giskard and DeepTeam. Columns are the raw model endpoint, the application entry point, multi-turn attacks, and the tool-result channel. garak is strong on the raw model endpoint, weak on the application entry point, partial on multi-turn, and none on the tool-result channel. Promptfoo is weak on the raw endpoint, strong on the application entry point, strong on multi-turn, and none on the tool-result channel. Giskard and DeepTeam follow the same pattern as Promptfoo. The whole fourth column reads none. What each tool can put a hostile string on Raw model endpoint Application entry point Multi-turn escalation Tool-result channel garak Strong Not its target Partial None Promptfoo Via a provider Strong Strong None Giskard Via a wrapper Strong Strong None DeepTeam Via a wrapper Strong Strong None Strong Partial Weak or absent by default
Three of the four columns are well served. The fourth is empty, and it is where agents get compromised.

The channel your agent is actually attacked through

Four boundaries where a hostile string can enter an agent's context An agent loop drawn in the centre. Four entry boundaries surround it. Boundary one, the user prompt, is covered by all four scanners. Boundary two, the system prompt and tool descriptions, is partly covered. Boundary three, the tool result returned by a search index, an API or a fetched page, is drawn as the accent block and is reached by none of the four by default; an arrow carries it back into the agent's context on the next turn. Boundary four, the rendered output, is covered only by output detectors. Where a hostile string can enter Agent loop model + context window + tool calls 1 · User prompt What a person types. covered by all four 2 · System prompt + tool docs Descriptions the model reads. partly covered 3 · Tool result Retrieved document, API field, fetched page, log line. reached by none of the four 4 · Rendered output What a downstream surface shows. output detectors only tool call out Scanners construct a string and hand it to boundary 1. An agent is compromised at boundary 3, where the text arrives as data it asked for.
Scanners stand at boundary one. Agent compromises arrive at boundary three.

An agent has at least four places a string can enter its context, and they are not equally defended. The user prompt is the one everybody tests. The system prompt and tool descriptions are a second surface — poisoned tool descriptions are a live, catalogued attack, covered in tool poisoning. The fourth is the rendering boundary on the way out.

The third is the interesting one. Every tool result an agent reads — a retrieved chunk, a JSON field from an API, the text of a page it was asked to summarise, a CI log, a WAF's own block record — is untrusted input that arrives wearing the costume of data the agent requested. That asymmetry is the whole of indirect prompt injection, and it is why the wiki keeps returning to it: prompt injection 101 for the shape, telemetry and logs as untrusted input for the channel people forget.

Now look at what a scanner does. It constructs a hostile string and hands it to a target through some interface. For garak the interface is a model generator — a completion call. For Promptfoo, Giskard and DeepTeam it is whatever callable you wrap: an HTTP endpoint, a Python function, a chat handler. In every case the string enters at boundary one. To reach boundary three you would have to stand up a hostile tool — a fake search index, a poisoned document store, an API that returns instructions in a description field — and none of the four ships that as a default mode.

This is not a bug in these projects. They were built to test LLM applications, and for a chatbot boundary one is the attack surface. It becomes a gap the moment your application calls tools, because then the model is reading text nobody typed.

The workaround is cheap and nobody does it: write one tool whose sole purpose is to return attacker-controlled content, register it in your agent's real tool catalogue, and drive the agent with ordinary benign prompts. Each of these four can then drive the harness — you are just using them as the runner rather than the payload source. The payloads come from their corpora; the delivery is yours.

The four, briefly

garak — the model scanner

NVIDIA's garak is the only one of the four whose unit of analysis is the model itself. Its architecture is three plugin families — probes that generate attacks, detectors that decide whether an attack landed, and generators that speak to a backend — with a probe catalogue in the low hundreds and a couple of dozen backends covering hosted APIs, Hugging Face and local runtimes. Point it at an endpoint, get a report.

That makes it excellent for a question the others answer badly: is this base model, before any of my prompting, prone to this class of failure? Useful when you are selecting a model, adding an open-weights option, or justifying a swap. Useless for telling you whether your RAG system leaks another tenant's documents, because garak has no idea your application exists.

Promptfoo — the CI generalist

Promptfoo's centre of gravity is a YAML file. You declare providers, test cases, assertions and red-team plugins; it generates attacks, runs them, grades the results and fails your build. It is the most natural of the four to wire into a pipeline, and the config-as-code shape means a finding becomes a regression test in one commit — the discipline that eval-driven development argues for everywhere else.

It is also now an OpenAI property. The acquisition was announced in March 2026 with a commitment to keep the open-source tooling open, and by the company's own reporting the project had reached six figures of developers and teams at a quarter of the Fortune 500 by then. Take the commitment at face value and still notice what it means structurally — more on that below.

Giskard — the scan with a review surface

Giskard is a Python library that scans an application for security and quality problems in the same pass: injection and leakage alongside hallucination, bias and robustness. It runs multi-turn attack agents rather than only single-shot probes, which matters because the interesting jailbreaks are conversational, and it pairs the library with a hub where non-engineers can look at findings.

The mixed taxonomy is the trade. If you want one run that tells a product owner where the model is unreliable and where it is unsafe, Giskard is the shortest path. If you need to hand an auditor a table whose rows line up with a named framework, the mixing is friction.

DeepTeam — the framework-mapped one

DeepTeam, from Confident AI, is built around an explicit two-axis model: a list of vulnerabilities (what can go wrong) crossed with a list of attack methods (how you try to make it), currently dozens of each, with the vulnerability list mapped onto the OWASP Top 10 for LLM Applications, NIST AI RMF and MITRE ATLAS. It is the youngest and smallest of the four by adoption, with a stable v1 and a following in the low thousands of stars.

The framework mapping is the reason to choose it. When the artefact you owe someone is evidence against a named control — and increasingly it is, per the compliance-regime comparison — a tool that emits results already keyed to the control saves a translation step that is otherwise done by hand, badly, at the end.

Probe count is the wrong axis

Every one of these projects leads with a number: probes, plugins, vulnerabilities, attack methods. The numbers are not comparable — garak's probe is a template, DeepTeam's pair is a cross-product cell, Promptfoo's plugin is a generator — and more importantly the number does not predict the thing you care about.

Two properties predict it better. The first is reach, above. The second is freshness, and it is the axis with no README badge: an attack corpus is a depreciating asset. Published jailbreaks get trained against. A probe set that was state of the art eighteen months ago now tests whether the vendor read the same papers you did, which is a real question but a much smaller one. A static corpus is a regression suite — genuinely valuable for catching a fix coming undone, and no longer a red team. That is the same lifecycle a private eval set goes through, and maintaining an eval set is the discipline that applies.

Which is why the maintenance question is not a footnote about project health. It is the buying criterion.

Who still maintains the corpus

Who maintains each attack corpus after the 2026 consolidation Three columns. PyRIT was archived by Microsoft on 27 March 2026 and is read-only, with no commits, releases or issue triage, yet is still widely recommended. Promptfoo was acquired by OpenAI in March 2026 with an open-source commitment, so the corpus testing your agent is curated by a model vendor. garak, Giskard and DeepTeam are maintained by an infrastructure vendor and two independent companies, and corpus freshness tracks each project's own release cadence. Corpus ownership after 2026 PyRIT — archived Read-only since 27 March 2026. No commits, releases or triage. Promptfoo — acquired OpenAI, March 2026, with an open-source commitment. garak · Giskard · DeepTeam One infrastructure vendor, two independents, each on its own release cadence. What it means for you What it means for you What it means for you A frozen reference, not a maintained tool — and still in most 2026 guides. A model vendor curates the threat model used against its competitors. Freshness is a project health question — check the last probe commit.
The stack consolidated in 2026. Two of the four leading corpora changed hands or stopped moving.

Two facts changed this landscape in 2026 and most published comparisons have not caught up.

Microsoft archived PyRIT on 27 March 2026. The repository is read-only: no commits, no releases, no issue triage. PyRIT was the reference implementation for multi-turn automated red teaming — Crescendo, TAP, skeleton-key style escalation — and it is still recommended by a large fraction of the "LLM red teaming in 2026" guides you will find this week. Whatever you pip-installed is the last version there will be. The techniques it encodes are still valid; the corpus around them has stopped tracking a field that moves monthly.

And Promptfoo is now OpenAI's, announced in March 2026 alongside a commitment to keep the open-source tools open. Assume that commitment holds — the structural point survives it. The corpus that decides whether your agent is safe is curated by a model vendor whose own agent products are in scope for the same attacks. There is no conspiracy required for that to matter: attention follows the acquirer's roadmap, and the failure classes that get probes are the ones that embarrass the acquirer's stack. If your deployment runs on a different vendor's models, you are now reading a threat model authored by a competitor of your supplier.

The practical consequence is that no single tool is a red team. Run one corpus you did not choose for its politics — garak against the model, one application scanner in CI — and keep a small local corpus of attacks that came from your own incidents and your own domain. That third set is the only one guaranteed to be about you, and it is usually twenty cases, not two thousand.

When to pick which

SituationReach forBecause
Choosing or swapping a base modelgarakIt is the only one that tests the model rather than your prompt around it.
Gating a build on security regressionsPromptfooYAML-in-repo turns a finding into a permanent test in one commit.
One report for both safety and qualityGiskardIts scan mixes injection and leakage with hallucination and bias by design.
Evidence keyed to OWASP, NIST or ATLASDeepTeamIts vulnerability list is already mapped; you skip a manual translation.
An agent that calls toolsAny of them, plus a hostile tool you writeNone of the four reaches the tool-result channel on its own.
Multi-turn escalation researchGiskard or DeepTeamPyRIT was the answer here and has been read-only since March 2026.

One thing worth saying plainly, because the tooling encourages the opposite: a scanner produces findings, not assurance. It cannot tell you what fraction of your agent's externally visible effects it was able to attack, and that fraction — not the pass rate — is what a reviewer should ask for. If your agent can send email, write to a database and open a pull request, and your red-team run only ever spoke to the chat endpoint, the report is about a different system than the one in production. Red-teaming agents covers what a real exercise adds on top, and evaluating guardrails and detectors covers the other half — whether the defence you deployed actually fires.

FAQ

Is PyRIT really unusable now?

It still runs, and its attack strategies are still sound. But the repository has been archived and read-only since 27 March 2026, so there are no new probes, no fixes and no issue triage. Treat it as a frozen reference implementation rather than a maintained tool.

Does OpenAI owning Promptfoo make it untrustworthy?

No. The open-source tooling remains open and the tool is good. The caution is narrower: an attack corpus reflects its maintainer's priorities, so do not let a vendor-curated corpus be the only one you run, particularly if you deploy on a different vendor's models.

Can I just run all four?

You can, and the overlap is large enough that the marginal finding drops fast. A better use of the same budget is one model-level scan, one application scan in CI, and the hostile-tool harness that none of them gives you.

Why does the tool-result channel matter more than the user prompt?

Because the user prompt is the channel you already distrust. Tool results arrive as data the agent asked for, are rarely re-checked, and frequently carry text written by someone else entirely — a web page, a ticket, a log line, a document in a shared drive.

What is the cheapest thing to do this week?

Write the hostile tool. Twenty lines that return a plausible document containing an instruction, registered in your real catalogue, driven by a benign prompt. Most teams find out in an afternoon whether their agent follows instructions it read rather than instructions it was given.

Further reading

On this wiki:

Project sources: