A clean red-team report on an agent usually means the scanner could not reach the place the attack arrives. All four of these tools send adversarial strings down the channel a user types into; an agent's real injection comes back inbound, inside a tool result it was told to trust. Choose on reach — which boundary the tool can place a hostile string on — and treat the probe count in the README as marketing.
At a glance
Four open-source projects, four different answers to "what is the thing under test".
| Tool | Maintainer | Unit of test | What it points at |
|---|---|---|---|
| garak | NVIDIA | A probe against a generator | A model endpoint. |
| Promptfoo | OpenAI (acquired 2026) | A YAML test case, graded | Your application's entry point. |
| Giskard | Giskard AI | A scan, plus multi-turn attacks | Your application, with a review surface. |
| DeepTeam | Confident AI | A vulnerability × attack pair | Your application, mapped to a framework. |
The channel your agent is actually attacked through
An agent has at least four places a string can enter its context, and they are not equally defended. The user prompt is the one everybody tests. The system prompt and tool descriptions are a second surface — poisoned tool descriptions are a live, catalogued attack, covered in tool poisoning. The fourth is the rendering boundary on the way out.
The third is the interesting one. Every tool result an agent reads — a retrieved chunk, a JSON field from an API, the text of a page it was asked to summarise, a CI log, a WAF's own block record — is untrusted input that arrives wearing the costume of data the agent requested. That asymmetry is the whole of indirect prompt injection, and it is why the wiki keeps returning to it: prompt injection 101 for the shape, telemetry and logs as untrusted input for the channel people forget.
Now look at what a scanner does. It constructs a hostile string and hands it to a target through some interface. For garak the interface is a model generator — a completion call. For Promptfoo, Giskard and DeepTeam it is whatever callable you wrap: an HTTP endpoint, a Python function, a chat handler. In every case the string enters at boundary one. To reach boundary three you would have to stand up a hostile tool — a fake search index, a poisoned document store, an API that returns instructions in a description field — and none of the four ships that as a default mode.
This is not a bug in these projects. They were built to test LLM applications, and for a chatbot boundary one is the attack surface. It becomes a gap the moment your application calls tools, because then the model is reading text nobody typed.
The workaround is cheap and nobody does it: write one tool whose sole purpose is to return attacker-controlled content, register it in your agent's real tool catalogue, and drive the agent with ordinary benign prompts. Each of these four can then drive the harness — you are just using them as the runner rather than the payload source. The payloads come from their corpora; the delivery is yours.
The four, briefly
garak — the model scanner
NVIDIA's garak is the only one of the four whose unit of analysis is the model itself. Its architecture is three plugin families — probes that generate attacks, detectors that decide whether an attack landed, and generators that speak to a backend — with a probe catalogue in the low hundreds and a couple of dozen backends covering hosted APIs, Hugging Face and local runtimes. Point it at an endpoint, get a report.
That makes it excellent for a question the others answer badly: is this base model, before any of my prompting, prone to this class of failure? Useful when you are selecting a model, adding an open-weights option, or justifying a swap. Useless for telling you whether your RAG system leaks another tenant's documents, because garak has no idea your application exists.
Promptfoo — the CI generalist
Promptfoo's centre of gravity is a YAML file. You declare providers, test cases, assertions and red-team plugins; it generates attacks, runs them, grades the results and fails your build. It is the most natural of the four to wire into a pipeline, and the config-as-code shape means a finding becomes a regression test in one commit — the discipline that eval-driven development argues for everywhere else.
It is also now an OpenAI property. The acquisition was announced in March 2026 with a commitment to keep the open-source tooling open, and by the company's own reporting the project had reached six figures of developers and teams at a quarter of the Fortune 500 by then. Take the commitment at face value and still notice what it means structurally — more on that below.
Giskard — the scan with a review surface
Giskard is a Python library that scans an application for security and quality problems in the same pass: injection and leakage alongside hallucination, bias and robustness. It runs multi-turn attack agents rather than only single-shot probes, which matters because the interesting jailbreaks are conversational, and it pairs the library with a hub where non-engineers can look at findings.
The mixed taxonomy is the trade. If you want one run that tells a product owner where the model is unreliable and where it is unsafe, Giskard is the shortest path. If you need to hand an auditor a table whose rows line up with a named framework, the mixing is friction.
DeepTeam — the framework-mapped one
DeepTeam, from Confident AI, is built around an explicit two-axis model: a list of vulnerabilities (what can go wrong) crossed with a list of attack methods (how you try to make it), currently dozens of each, with the vulnerability list mapped onto the OWASP Top 10 for LLM Applications, NIST AI RMF and MITRE ATLAS. It is the youngest and smallest of the four by adoption, with a stable v1 and a following in the low thousands of stars.
The framework mapping is the reason to choose it. When the artefact you owe someone is evidence against a named control — and increasingly it is, per the compliance-regime comparison — a tool that emits results already keyed to the control saves a translation step that is otherwise done by hand, badly, at the end.
Probe count is the wrong axis
Every one of these projects leads with a number: probes, plugins, vulnerabilities, attack methods. The numbers are not comparable — garak's probe is a template, DeepTeam's pair is a cross-product cell, Promptfoo's plugin is a generator — and more importantly the number does not predict the thing you care about.
Two properties predict it better. The first is reach, above. The second is freshness, and it is the axis with no README badge: an attack corpus is a depreciating asset. Published jailbreaks get trained against. A probe set that was state of the art eighteen months ago now tests whether the vendor read the same papers you did, which is a real question but a much smaller one. A static corpus is a regression suite — genuinely valuable for catching a fix coming undone, and no longer a red team. That is the same lifecycle a private eval set goes through, and maintaining an eval set is the discipline that applies.
Which is why the maintenance question is not a footnote about project health. It is the buying criterion.
Who still maintains the corpus
Two facts changed this landscape in 2026 and most published comparisons have not caught up.
Microsoft archived PyRIT on 27 March 2026. The repository is read-only: no commits, no releases, no issue triage. PyRIT was the reference implementation for multi-turn automated red teaming — Crescendo, TAP, skeleton-key style escalation — and it is still recommended by a large fraction of the "LLM red teaming in 2026" guides you will find this week. Whatever you pip-installed is the last version there will be. The techniques it encodes are still valid; the corpus around them has stopped tracking a field that moves monthly.
And Promptfoo is now OpenAI's, announced in March 2026 alongside a commitment to keep the open-source tools open. Assume that commitment holds — the structural point survives it. The corpus that decides whether your agent is safe is curated by a model vendor whose own agent products are in scope for the same attacks. There is no conspiracy required for that to matter: attention follows the acquirer's roadmap, and the failure classes that get probes are the ones that embarrass the acquirer's stack. If your deployment runs on a different vendor's models, you are now reading a threat model authored by a competitor of your supplier.
The practical consequence is that no single tool is a red team. Run one corpus you did not choose for its politics — garak against the model, one application scanner in CI — and keep a small local corpus of attacks that came from your own incidents and your own domain. That third set is the only one guaranteed to be about you, and it is usually twenty cases, not two thousand.
When to pick which
| Situation | Reach for | Because |
|---|---|---|
| Choosing or swapping a base model | garak | It is the only one that tests the model rather than your prompt around it. |
| Gating a build on security regressions | Promptfoo | YAML-in-repo turns a finding into a permanent test in one commit. |
| One report for both safety and quality | Giskard | Its scan mixes injection and leakage with hallucination and bias by design. |
| Evidence keyed to OWASP, NIST or ATLAS | DeepTeam | Its vulnerability list is already mapped; you skip a manual translation. |
| An agent that calls tools | Any of them, plus a hostile tool you write | None of the four reaches the tool-result channel on its own. |
| Multi-turn escalation research | Giskard or DeepTeam | PyRIT was the answer here and has been read-only since March 2026. |
One thing worth saying plainly, because the tooling encourages the opposite: a scanner produces findings, not assurance. It cannot tell you what fraction of your agent's externally visible effects it was able to attack, and that fraction — not the pass rate — is what a reviewer should ask for. If your agent can send email, write to a database and open a pull request, and your red-team run only ever spoke to the chat endpoint, the report is about a different system than the one in production. Red-teaming agents covers what a real exercise adds on top, and evaluating guardrails and detectors covers the other half — whether the defence you deployed actually fires.
FAQ
Is PyRIT really unusable now?
It still runs, and its attack strategies are still sound. But the repository has been archived and read-only since 27 March 2026, so there are no new probes, no fixes and no issue triage. Treat it as a frozen reference implementation rather than a maintained tool.
Does OpenAI owning Promptfoo make it untrustworthy?
No. The open-source tooling remains open and the tool is good. The caution is narrower: an attack corpus reflects its maintainer's priorities, so do not let a vendor-curated corpus be the only one you run, particularly if you deploy on a different vendor's models.
Can I just run all four?
You can, and the overlap is large enough that the marginal finding drops fast. A better use of the same budget is one model-level scan, one application scan in CI, and the hostile-tool harness that none of them gives you.
Why does the tool-result channel matter more than the user prompt?
Because the user prompt is the channel you already distrust. Tool results arrive as data the agent asked for, are rarely re-checked, and frequently carry text written by someone else entirely — a web page, a ticket, a log line, a document in a shared drive.
What is the cheapest thing to do this week?
Write the hostile tool. Twenty lines that return a plausible document containing an instruction, registered in your real catalogue, driven by a benign prompt. Most teams find out in an afternoon whether their agent follows instructions it read rather than instructions it was given.
Further reading
On this wiki:
- Red-teaming agents — what a structured exercise adds beyond a scanner run.
- Red-teaming & safety evaluation — the operational cadence.
- Prompt injection 101 — the direct and indirect shapes.
- Tool poisoning — injection via tool descriptions.
- Evaluating guardrails & detectors — measuring the defence, not the attack.
- Four shapes of a guardrail — the runtime-defence counterpart to this comparison.
Project sources:
- NVIDIA/garak — probes, detectors, generators.
- Promptfoo is joining OpenAI — the acquisition announcement.
- Giskard — AI red teaming field guide.
- confident-ai/deepteam — vulnerabilities and attack methods.
- Azure/PyRIT — archived 27 March 2026.