AI Blog

The scaffold found the bug, not the model

A startup’s analyzer took six CVEs out of curl in a window where, by its own account, Codex and Mythos found none — and a 2026 benchmark recovers 68% of real AI-found CVEs using only small and open-weight models, with no frontier model in the detection path. The variable that moved is the search structure, not the model. The number to buy on is accepted findings per maintainer-hour: 29 reports were filed and six were accepted, all rated Low.

By Agentic AI Wiki 11 min read

A startup's analyzer found six CVEs in curl — one of the most audited C codebases on earth — in a window where, by its own account, OpenAI's Codex and Anthropic's Mythos found none. The tempting read is that somebody has a better model. The published evidence says the opposite: a benchmark released this year rediscovers 68% of real AI-found CVEs using only small and open-weight models, and no frontier model performs detection anywhere in it. The variable that moved was the scaffold, and the number you should be shopping on is not recall.

At a glance

Two data points from 2026, one commercial and one academic, pointing the same way.

EventWhat was claimedWhat it rests onWhy it matters
curl 8.22.0 (2 Sep 2026) Six of nine advisories credit one AISLE reporter; all rated Low curl's own advisories; AISLE's disclosure post A small vendor out-finds two frontier labs on a hardened target
curl 8.21.0 (Jun 2026) Eighteen CVEs, a record for one curl release; AISLE credited with six, including the project's oldest issue curl release advisories The finding rate is not a one-off
HoF-Bench (arXiv 2607.27030) A minimal analyzer rediscovers up to 65 of 95 real CVEs; no frontier model used for detection 95 public AI-found CVEs across eight repos, pinned at vulnerable commits Isolates the scaffold from the model, and the scaffold wins
curl's bug bounty (Jan 2026) Closed on HackerOne after a flood of fabricated AI reports Maintainer's public statement The constraint is precision, not recall
Reports filed versus CVEs accepted Horizontal bar chart. Twenty-nine reports were filed against curl in late August 2026; six were accepted as CVEs; none were rated above Low severity. The gap between the first two bars is triage work carried by maintainers. One disclosure run against curl, August–September 2026 Reports filed 29 Accepted as CVEs 6 Rated above Low 0 10 20 30 0 reports
AISLE says it filed 29 reports; six became CVEs. The other 23 are somebody's afternoon.

What actually happened in curl

curl 8.22.0 shipped on 2 September 2026 with nine security advisories. Six of them credit the same reporter — Stanislav Fort of AISLE — for findings filed over four days in late August: a use-after-free in the OpenSSL provider path, a pinning bypass, a native CA store connection-reuse issue, a secure-cookie attribute bypass involving a tab character, a wolfSSL CA-cache callback override, and a domain-scoped public-suffix cookie problem. All six are rated Low.

It was not the first batch. AISLE is also credited with six of the eighteen CVEs in curl 8.21.0 in June — a record count for a single curl release — including CVE-2026-8932, which first shipped in curl 7.7 on 22 March 2001 and is the oldest security issue in the project's history. Twenty-five years of fuzzing, static analysis, paid audits and a bug bounty went past it.

AISLE's claim that Codex and Mythos found nothing over the same window is the vendor's own, and should be read as such. But the part that does not depend on trusting a vendor is the shape of the artefact: a small company with no frontier model of its own, pointing a purpose-built analyzer at a codebase that everyone has already looked at, and getting acceptances.

The benchmark that separates the two variables

The fixed detection scaffold A repository pinned at its vulnerable commit, with target-file scope but no CVE metadata, feeds four repeated detection passes around a small model. An optional context-generation stage and a multi-round triage stage follow, and a detector-blinded judge accepts a finding only when code path, root cause, attack condition and impact all match. Everything outside the model box is under the experimenter's control Repo pinned at the vulnerable commit + target-file scope only Context generation optional stage Detection — four repeated passes Small / open-weight 3–13B active Candidate findings deduplicated pass 2, 3, 4 re-run with variation Multi-round triage replayable; drops what cannot be substantiated Blinded judge no CVE metadata shown Accepted only if all four match same code path · same root cause same attack condition · same impact 65 of 95 CVEs recovered no frontier model in the detection path
The fixed scaffold in HoF-Bench. Everything outside the model box is the part under the experimenter's control.

When a vendor beats a frontier model, the useful question is which variable moved. HoF-Bench is built to answer it. The benchmark takes 95 public AI-discovered CVEs across eight repositories, pins each at its vulnerable commit, and gives the analyzer the source and a target-file scope — but not the CVE identifier, its description, the fix, or the expected mechanism. A detector-blinded judge then credits a finding only if it identifies the same code path, the same root cause, the same attack condition and the same impact. That is a much harsher bar than "did it mention a buffer".

The result

A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 — about 68% — under that protocol. The ten detector backbones are five open-weight models (21B–284B total parameters, 3–13B active) and five proprietary small or flash-tier models. No frontier model performs detection anywhere in the study. All detectors run inside the same fixed scaffold: four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage, across 7,600 model-CVE pass records.

What that implies

If two-thirds of real, already-found vulnerabilities are recoverable by a 3–13B active-parameter model repeated four times inside a decent harness, then detection capability was not the binding constraint for those bugs. Search structure was. This is the concrete version of an argument that is usually made qualitatively: every agentic score is a joint score on a model and the harness it ran inside, and only one of the two gets named in the headline.

Note also where the misses concentrate. Difficulty in HoF-Bench is strongly structured by language, and the CVEs that every model misses cluster in C infrastructure code — which is, inconveniently, exactly the category curl belongs to. The remaining third is not evenly distributed noise; it is a specific kind of bug that a repeated cheap pass does not reach.

Recall is getting cheap. Precision is the thing nobody is selling

Now the number that should change your plans. AISLE says it filed 29 reports; six were accepted. That is roughly a one-in-five acceptance rate, on a good run, by a specialist, against a project whose maintainers are unusually rigorous — and every one of the other 23 consumed a human's attention.

curl closed its HackerOne bug bounty in January 2026 for exactly this reason: a flood of confident, detailed, fabricated reports from people pointing LLMs at the codebase for bounty money. The project that then accepted twelve AISLE CVEs across two releases is the same project that had already shut the front door, because the difference between a good analyzer and a slop generator is invisible until someone reads the report.

Three numbers that describe an automated vulnerability finder Three columns — recall, precision and severity — each with who pays the cost and how fast it is improving. Recall is falling in price, precision is paid by maintainers, and severity decides whether the output is hardening or averted incidents. Buy on the second column; the first is already commoditised Recall Precision Severity 68% of real CVEs recovered by small models, four passes 6 accepted from 29 filed on a specialist's good run 6 of 6 rated Low hardening, not averted breach Paid by: whoever rents the GPU Trend: falling every quarter Paid by: the maintainer reading it Trend: not improving on its own Paid by: the team that must patch Trend: set by the target's maturity The metric to put on the contract: accepted findings per maintainer-hour of triage Findings per repository is a vanity number, and it is the one every vendor reports.
Three numbers describe an automated finder. Two of them are paid for by someone who did not buy it.

Severity is the third number and it is the quiet one. All six of the September findings are rated Low. That is not a criticism — a real Low in curl is a real bug — but it does mean the value of the batch is measured in hardening, not in averted incidents, and hardening competes for the same maintainer hours as everything else. An automated finder that produces twenty Lows for every accepted one is not obviously a gift.

What to do differently if you are buying or building one

The procurement instinct is to ask which model a security product uses. On this evidence that is the least informative question available.

  • Ask about the scaffold, in detail. How many passes, with what variation between them. Whether there is a context-generation stage and what it feeds on. How triage works and how many rounds it runs. Whether findings are de-duplicated across passes before a human sees them. These are the parts that moved the number in the only controlled comparison available.
  • Make acceptance the metric, not findings. Findings per repository is a vanity number. Accepted findings per maintainer-hour of triage is the number that decides whether the tool is net positive, and it is measurable on your own backlog in a week.
  • Budget for triage before you budget for tokens. A cheap detector with a bad accept rate transfers its cost onto the people you were trying to help, and unlike inference that cost does not fall every quarter.
  • Evaluate on your own language and your own code. The C infrastructure result says difficulty is structured by what you wrote, not by what the leaderboard averaged. Pin ten of your own historical CVEs at their vulnerable commits and run the tool blind — that is a two-day build and it beats every vendor benchmark.
  • Separate finding from fixing. These are different capabilities with different failure costs; see vulnerability remediation agents for the second half.

The part that generalises past security

The specific claim here is about vulnerability discovery. The general one is about any task where the answer can be checked more cheaply than it can be produced. Detection is the purest example: a candidate finding costs one pass, and confirming it costs a triage round — so repetition pays, and repetition with a cheap model pays more than one attempt with an expensive one.

That is the generator–verifier gap working in your favour, and it inverts the usual advice. Where verification is cheap, buy passes. Where verification is expensive — most agent work, where checking the output means doing the work again — buy capability. Most teams apply one rule to both cases, and the curl result is a reminder of what it costs to get the classification wrong in the direction of the expensive model.

FAQ

Does this mean frontier models are bad at finding vulnerabilities?

No. It means that in the one study that controls for the scaffold, frontier models were not needed to recover two-thirds of real CVEs, and that a specialist vendor's scaffold beat general-purpose coding agents on a specific hardened target. Both are claims about configuration, not about ceilings.

Is AISLE's comparison against Codex and Mythos independently verified?

Not as far as we can tell. The CVE credits in curl's advisories are public and checkable; the claim about what the other two systems found over the same window is the vendor's own and rests on their account of how they were run.

Why are all six findings rated Low?

curl rates severity conservatively and these are edge conditions in TLS backends and cookie handling rather than remote code execution. It is a fair signal about what this class of tooling currently surfaces on a mature target: many real, narrow bugs rather than a small number of critical ones.

Should we point a coding agent at our own repository this week?

You can, but decide first who reads the output. Run it, triage everything it produces yourself, and measure the accept rate before anyone else sees a report. A team that ships unverified agent findings to another team is recreating the problem that closed curl's bounty.

Does a higher accept rate just mean the tool is being conservative?

Sometimes, which is why you track both numbers on the same backlog. Accept rate alone can be gamed by reporting only the obvious; pair it with rediscovery rate on your own historical CVEs, pinned at their vulnerable commits, and gaming either one costs the other.

Further reading

On this wiki:

Sources: