Google's threat intelligence group published a vulnerability-trends report on 30 September 2026 and the line everyone quoted — exactly 50% of AI-discovered flaws yield remote code execution, against 26% of everything else — is the one you can act on least, because GTIG publishes no sample size for it and its own attribution method selects for the handful of vendors currently pointing agents at OpenSSL and the kernel. Buried four sections down is a number with your name on it: 782 CVEs in agent frameworks and orchestration in eight months, nearly nine times the count for frontier models themselves. The risk that moved this year is in your dependency tree, not in the model.
At a glance
Four numbers from the same report, and what each one can actually carry.
| Figure | Value | How solid | What it supports |
|---|---|---|---|
| RCE share, AI-discovered | 50% vs 26% | Verbatim, no denominator published | A hypothesis about the sample, not a planning input |
| Monthly disclosures | 5,045 → 10,740 | Jan to Aug 2026, GTIG's own dataset | Triage load doubled; independently corroborated |
| Time to exploitation | 4 days, one CVE | Single case, same day as public PoC | Roughly the pre-AI 2023 average — not acceleration |
| Orchestration-layer CVEs | 782 in 8 months | Counted, layer-attributed, uncontested | A patch SLA you probably do not have |
What the report says
Vulnerability Discovery and Exploitation Trends in the AI Era, by Robin Grunewald, Supriya Mazumdar and Kelli Vanderlee, covers disclosures from 1 January 2025 to 31 August 2026. Its headline findings are consistent and, where they can be cross-checked, broadly right.
Monthly disclosures doubled inside 2026 — 5,045 in January, 10,477 in July, 10,740 in August. Independent counts from the NVD feed run higher still (12,291 for August), which is what you would expect from a narrower curated dataset, so the direction is not in dispute. Exploitation rose too, from an average of 10.5 vulnerabilities per month in 2025 to 18 across January–August 2026, with zero-day exploitation up more modestly from 8 to 11. And 141 distinct vulnerabilities were both disclosed and exploited in those eight months, beating all of 2025's 127 — while amounting to 0.23% of disclosures, roughly one in 431.
GTIG is careful about its own baselines in a way the coverage was not. It notes that automated CNA assignment inflates the denominator — descriptions containing "Linux Kernel" alone generated about 5,000 CVEs between January and August 2026 with no observed in-the-wild zero-days — and that its risk ratings are GTIG's own, not CVSS severity. Both caveats cut against the simplest reading of the trend.
The 50% describes who is running agents, not what agents find
Read GTIG's attribution method and the finding changes character. A vulnerability enters the AI-discovered set two ways: by ingesting confirmed disclosures from frontier AI research programmes, or by programmatically parsing CISA advisories, MITRE records and vendor bulletins for an explicit acknowledgment that an autonomous agent found it. Both filters select on disclosure practice. Neither observes discovery.
So the sample is whoever currently says so in an advisory — and in 2026 that is a short list. GTIG names Hacktron AI and AISLE as its anchors. AISLE became the first AI-native CVE Numbering Authority in July 2026 and has published runs of OpenSSL, curl and Wireshark findings; Hacktron's public work is variant analysis against C and C++ network-facing software. Those targets are memory-unsafe systems code. Memory-unsafe systems code yields memory corruption, and memory corruption yields code execution. The target mix predicts the result without any claim about agent capability at all.
GTIG's own explanation is hedged — the concentration "likely stems from" agents navigating multi-step semantic code paths across C/C++ libraries, runtimes and hypervisors. That is a capability story for the same observation, and it may well be right. The problem is that the report publishes no N for the AI-discovered column and no control for target-software composition, so there is no way to tell the two apart from the outside. "Exactly 50%" is itself a hint: round halves come from small denominators.
There is also contrary evidence from the same year. VulnCheck's State of Exploitation 1H-2026, published in late July, attributed 1,061 vulnerabilities to AI-assisted discovery and found 14 of them — 1.3% — confirmed exploited in the wild, which its author called roughly the same as every other vulnerability in the period and below the historical average. Anthropic's Project Glasswing is cited there with more than 23,000 findings, 126 published CVEs and one confirmed exploited.
These two results are not actually in conflict, and it is worth not staging a fight. GTIG measured potential impact class; VulnCheck measured realised exploitation. AI-found bugs can be more severe on paper and no more likely to be attacked. GTIG concedes as much: confirmed exploitation of AI-discovered flaws is "an early indicator rather than an established trend".
"Within four days" is approximately the 2023 norm
The report's worked case is CVE-2026-1731, an unauthenticated OS command injection in BeyondTrust Privileged Remote Access and Remote Support, found autonomously by Hacktron's research agent. GTIG observed one threat cluster exploiting it within four days of public disclosure and five more within seven, with post-exploitation including privilege escalation, data exfiltration, and SNOWLIGHT, SPARKRAT and cryptominer payloads.
Four days sounds like the future arriving. Put it against Mandiant's own time-to-exploit series and it is the recent past: average TTE fell from 63 days across 2018–19 to 44, then 32, then 5 days for the 2023 cohort — and within that cohort 12% of n-days were exploited inside one day and 29% inside a week. A four-day n-day in 2026 sits at the 2023 average and is slower than the fastest eighth of 2023.
The proximate accelerant also was not AI. Exploitation began the day a public proof-of-concept was posted. The four-day figure is measured from disclosure, but what it really records is a PoC-to-exploitation interval of about zero — the oldest pattern in vulnerability management. Attributing the speed to the flaw's AI provenance is a causal step GTIG does not take, and neither should you.
The number nobody quoted
Now the part of the report that is about the systems this wiki is about. GTIG counted 2,076 cumulative CVEs in the AI stack from January 2025 to August 2026, of which more than 1,500 landed in the first eight months of 2026 alone. Broken out by layer for 2026: AI orchestration and agent frameworks 782, AI web apps and portals 230, inference and serving infrastructure 212, model security advisories 106, ML frameworks and hubs 99, frontier models 97, MLOps and experiment tracking 39, vector databases and search 19.
Hold those first and sixth numbers next to each other. The orchestration layer produced roughly eight times as many CVEs as the frontier models did. That inverts where attention goes: the model is the thing with the system card, the safety evaluations, the red-team report and the vendor's security team. The orchestration layer is the thing that arrived as a pip install during a prototype, holds the credentials, owns the tool catalogue and sits on the egress path. It is the layer with the largest blast radius and the least patch discipline, and it is now the fastest-growing CVE category in the stack.
The honest caveat is that volume is not risk. Counting CVEs in a category partly counts how many CNAs are active in it, and a young ecosystem shipping fast with new scanner attention will mint CVEs quickly. GTIG also reports that zero-day exploitation of AI infrastructure has not been observed at all, and that of those 2,076 disclosures only a handful are confirmed exploited in the wild. So this is not a claim that your framework is under attack. It is a claim about where your patch work is going to come from, and it has already started.
What to change this week
None of this requires a view on whether agents find scarier bugs. It requires knowing what you depend on and how fast you can replace it.
- Produce the orchestration-layer inventory. The framework, every MCP server, the gateway, the sandbox runtime, the vector store, the browser automation. Version-pinned, with an owner each. Most teams can name the framework and nothing else — see agent inventory and registry for deriving it rather than asking for it.
- Set a patch SLA for that layer specifically, and measure against it. At 782 CVEs in eight months the arrival rate for this layer alone is a couple a day across the ecosystem; a quarterly upgrade cadence is not a policy, it is a backlog. Vulnerability management for agent platforms covers what the SLA has to cover.
- Stop treating "unauthenticated" as the only severity that matters. The BeyondTrust flaw was pre-auth, which is why it got attention. The ones in your agent stack will mostly be post-auth, and post-auth is where an agent lives — it is already inside, holding a token, with ambient reach. Blast radius is the frame that survives this.
- Subscribe to the advisories for the things you install, not just the things you buy. An MCP server from a GitHub account has no advisory feed and no deprecation policy. That is a procurement decision you already made — third-party model and vendor risk, and agent supply-chain security for the install path.
- Pin by digest and check the digest. Doubling disclosure volume means more frequent forced upgrades, and more upgrades means more chances for the upgrade itself to be the compromise. Pinning and verification.
If you only do one of these, do the inventory. Everything else is unschedulable without it, and the inventory is the artefact that turns a report like this one from a headline into a work queue.
FAQ
Is the 50% RCE figure wrong?
No — it is quoted accurately from GTIG and there is no reason to doubt the arithmetic on their sample. The question is what the sample represents. Because membership is determined by whether an advisory explicitly credits an AI agent, the set is dominated by a few prolific vendors working on memory-unsafe systems code, which predicts an RCE-heavy mix by itself.
Should I reprioritise my patch queue because of this report?
Not on the basis of the AI-provenance signal, which you cannot even observe — public CVE records have no AI-attribution field, as GTIG itself notes. Reprioritise on the layer finding instead: if your orchestration and agent-framework dependencies are not in your vulnerability-management scope, that is a concrete gap and fixing it does not depend on any contested claim.
Does a doubling of disclosures mean twice the risk?
No. GTIG puts the disclosed-and-exploited rate at 0.23%, about one in 431, and attributes part of the volume growth to automated CNA assignment — around 5,000 Linux-kernel CVEs in eight months with no observed in-the-wild zero-days. What doubled reliably is triage cost.
Why so many CVEs in agent frameworks and so few in frontier models?
Partly real exposure — orchestration code is network-facing, handles untrusted input and holds credentials — and partly accounting. Model issues are frequently handled as advisories or silently patched in a hosted service without a CVE ID, which GTIG lists as one reason public data undercounts. Read 782 against 97 as a statement about where patchable, trackable defects live, not as a safety comparison.
What is the one measurement worth adding?
Mean age of your orchestration-layer dependencies against their latest release, reported per deployed agent. It is cheap, it moves when you do the work, and it is the number that would have told you in advance whether a doubling of disclosure volume was going to be survivable.
Further reading
On this wiki:
- Vulnerability Management for Agent Platforms — the patch SLA this report argues for.
- Agent Inventory & Registry — deriving the dependency list instead of requesting it.
- Agent Supply-Chain Security — how the orchestration layer gets installed in the first place.
- Vulnerability Remediation Agents — the other side of the volume problem.
- Unpinned Vendor Defaults — why your stack changes without a deploy.
Sources:
- GTIG — Vulnerability Discovery and Exploitation Trends in the AI Era, 30 September 2026.
- VulnCheck — State of Exploitation 1H-2026, the 1.3% exploited figure on 1,061 AI-attributed vulnerabilities.
- Mandiant — 2023 time-to-exploit trends, the five-day average and the n-day distribution.
- Mandiant — 2021–2022 time-to-exploit trends, the 63/44/32-day series.