A startup's analyzer found six CVEs in curl — one of the most audited C codebases on earth — in a window where, by its own account, OpenAI's Codex and Anthropic's Mythos found none. The tempting read is that somebody has a better model. The published evidence says the opposite: a benchmark released this year rediscovers 68% of real AI-found CVEs using only small and open-weight models, and no frontier model performs detection anywhere in it. The variable that moved was the scaffold, and the number you should be shopping on is not recall.
At a glance
Two data points from 2026, one commercial and one academic, pointing the same way.
| Event | What was claimed | What it rests on | Why it matters |
|---|---|---|---|
| curl 8.22.0 (2 Sep 2026) | Six of nine advisories credit one AISLE reporter; all rated Low | curl's own advisories; AISLE's disclosure post | A small vendor out-finds two frontier labs on a hardened target |
| curl 8.21.0 (Jun 2026) | Eighteen CVEs, a record for one curl release; AISLE credited with six, including the project's oldest issue | curl release advisories | The finding rate is not a one-off |
| HoF-Bench (arXiv 2607.27030) | A minimal analyzer rediscovers up to 65 of 95 real CVEs; no frontier model used for detection | 95 public AI-found CVEs across eight repos, pinned at vulnerable commits | Isolates the scaffold from the model, and the scaffold wins |
| curl's bug bounty (Jan 2026) | Closed on HackerOne after a flood of fabricated AI reports | Maintainer's public statement | The constraint is precision, not recall |
What actually happened in curl
curl 8.22.0 shipped on 2 September 2026 with nine security advisories. Six of them credit the same reporter — Stanislav Fort of AISLE — for findings filed over four days in late August: a use-after-free in the OpenSSL provider path, a pinning bypass, a native CA store connection-reuse issue, a secure-cookie attribute bypass involving a tab character, a wolfSSL CA-cache callback override, and a domain-scoped public-suffix cookie problem. All six are rated Low.
It was not the first batch. AISLE is also credited with six of the eighteen CVEs in curl 8.21.0 in June — a record count for a single curl release — including CVE-2026-8932, which first shipped in curl 7.7 on 22 March 2001 and is the oldest security issue in the project's history. Twenty-five years of fuzzing, static analysis, paid audits and a bug bounty went past it.
AISLE's claim that Codex and Mythos found nothing over the same window is the vendor's own, and should be read as such. But the part that does not depend on trusting a vendor is the shape of the artefact: a small company with no frontier model of its own, pointing a purpose-built analyzer at a codebase that everyone has already looked at, and getting acceptances.
The benchmark that separates the two variables
When a vendor beats a frontier model, the useful question is which variable moved. HoF-Bench is built to answer it. The benchmark takes 95 public AI-discovered CVEs across eight repositories, pins each at its vulnerable commit, and gives the analyzer the source and a target-file scope — but not the CVE identifier, its description, the fix, or the expected mechanism. A detector-blinded judge then credits a finding only if it identifies the same code path, the same root cause, the same attack condition and the same impact. That is a much harsher bar than "did it mention a buffer".
The result
A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 — about 68% — under that protocol. The ten detector backbones are five open-weight models (21B–284B total parameters, 3–13B active) and five proprietary small or flash-tier models. No frontier model performs detection anywhere in the study. All detectors run inside the same fixed scaffold: four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage, across 7,600 model-CVE pass records.
What that implies
If two-thirds of real, already-found vulnerabilities are recoverable by a 3–13B active-parameter model repeated four times inside a decent harness, then detection capability was not the binding constraint for those bugs. Search structure was. This is the concrete version of an argument that is usually made qualitatively: every agentic score is a joint score on a model and the harness it ran inside, and only one of the two gets named in the headline.
Note also where the misses concentrate. Difficulty in HoF-Bench is strongly structured by language, and the CVEs that every model misses cluster in C infrastructure code — which is, inconveniently, exactly the category curl belongs to. The remaining third is not evenly distributed noise; it is a specific kind of bug that a repeated cheap pass does not reach.
Recall is getting cheap. Precision is the thing nobody is selling
Now the number that should change your plans. AISLE says it filed 29 reports; six were accepted. That is roughly a one-in-five acceptance rate, on a good run, by a specialist, against a project whose maintainers are unusually rigorous — and every one of the other 23 consumed a human's attention.
curl closed its HackerOne bug bounty in January 2026 for exactly this reason: a flood of confident, detailed, fabricated reports from people pointing LLMs at the codebase for bounty money. The project that then accepted twelve AISLE CVEs across two releases is the same project that had already shut the front door, because the difference between a good analyzer and a slop generator is invisible until someone reads the report.
Severity is the third number and it is the quiet one. All six of the September findings are rated Low. That is not a criticism — a real Low in curl is a real bug — but it does mean the value of the batch is measured in hardening, not in averted incidents, and hardening competes for the same maintainer hours as everything else. An automated finder that produces twenty Lows for every accepted one is not obviously a gift.
What to do differently if you are buying or building one
The procurement instinct is to ask which model a security product uses. On this evidence that is the least informative question available.
- Ask about the scaffold, in detail. How many passes, with what variation between them. Whether there is a context-generation stage and what it feeds on. How triage works and how many rounds it runs. Whether findings are de-duplicated across passes before a human sees them. These are the parts that moved the number in the only controlled comparison available.
- Make acceptance the metric, not findings. Findings per repository is a vanity number. Accepted findings per maintainer-hour of triage is the number that decides whether the tool is net positive, and it is measurable on your own backlog in a week.
- Budget for triage before you budget for tokens. A cheap detector with a bad accept rate transfers its cost onto the people you were trying to help, and unlike inference that cost does not fall every quarter.
- Evaluate on your own language and your own code. The C infrastructure result says difficulty is structured by what you wrote, not by what the leaderboard averaged. Pin ten of your own historical CVEs at their vulnerable commits and run the tool blind — that is a two-day build and it beats every vendor benchmark.
- Separate finding from fixing. These are different capabilities with different failure costs; see vulnerability remediation agents for the second half.
The part that generalises past security
The specific claim here is about vulnerability discovery. The general one is about any task where the answer can be checked more cheaply than it can be produced. Detection is the purest example: a candidate finding costs one pass, and confirming it costs a triage round — so repetition pays, and repetition with a cheap model pays more than one attempt with an expensive one.
That is the generator–verifier gap working in your favour, and it inverts the usual advice. Where verification is cheap, buy passes. Where verification is expensive — most agent work, where checking the output means doing the work again — buy capability. Most teams apply one rule to both cases, and the curl result is a reminder of what it costs to get the classification wrong in the direction of the expensive model.
FAQ
Does this mean frontier models are bad at finding vulnerabilities?
No. It means that in the one study that controls for the scaffold, frontier models were not needed to recover two-thirds of real CVEs, and that a specialist vendor's scaffold beat general-purpose coding agents on a specific hardened target. Both are claims about configuration, not about ceilings.
Is AISLE's comparison against Codex and Mythos independently verified?
Not as far as we can tell. The CVE credits in curl's advisories are public and checkable; the claim about what the other two systems found over the same window is the vendor's own and rests on their account of how they were run.
Why are all six findings rated Low?
curl rates severity conservatively and these are edge conditions in TLS backends and cookie handling rather than remote code execution. It is a fair signal about what this class of tooling currently surfaces on a mature target: many real, narrow bugs rather than a small number of critical ones.
Should we point a coding agent at our own repository this week?
You can, but decide first who reads the output. Run it, triage everything it produces yourself, and measure the accept rate before anyone else sees a report. A team that ships unverified agent findings to another team is recreating the problem that closed curl's bounty.
Does a higher accept rate just mean the tool is being conservative?
Sometimes, which is why you track both numbers on the same backlog. Accept rate alone can be gamed by reporting only the obvious; pair it with rediscovery rate on your own historical CVEs, pinned at their vulnerable commits, and gaming either one costs the other.
Further reading
On this wiki:
- The Agent Harness — why every agentic benchmark number scores two things and names one.
- The Generator–Verifier Gap — when repetition beats capability, and when it does not.
- Agent Search Strategies — passes, breadth and re-sampling as a design choice.
- Vulnerability Remediation Agents — the fixing half of the problem.
- Benchmark Contamination — why pinning at the vulnerable commit matters.
Sources:
- curl security advisories — the CVE credits and severity ratings.
- HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models (arXiv 2607.27030).
- AISLE: six new CVEs in curl, including the oldest issue ever reported.
- curl 8.22.0 release notes.