A wrong severity score gets argued about in a meeting. A wrong applicability string produces no meeting at all, because the scanner returns no rows and nobody learns that a question was asked. NIST presented its AI agent enrichment workflow for the National Vulnerability Database today, and the interesting question is not whether the model is accurate — it is that enrichment emits three fields with three incompatible failure shapes, and the record format has nowhere to write down that a machine produced one of them.
At a glance
Five dates, one of which is a deadline you can still act on.
| Date | What happened | Why it matters |
|---|---|---|
| 15 April 2026 | NVD moves from universal to risk-based enrichment | Roughly 29,000 CVEs published before 1 March 2026 reclassified “Not Scheduled” |
| 12 August 2026 | RFI on modernising the NVD in the age of AI (docket NIST-2026-0100) | Asks which tasks suit AI, which need human review, what safeguards, how decisions stay auditable |
| 12 August 2026 | NIST discloses V-etalon | An AI-enabled tool to “aid in enriching vulnerability information”; unreleased, production role unstated |
| 17 September 2026 | ITL webinar: the agent enrichment workflow at the NVD | Architecture, problems found during implementation, early results |
| 13 October 2026 | RFI comments close, 11:59pm ET | Filed publicly, without redaction, on regulations.gov |
What NIST actually said
On 12 August NIST published a Request for Information on modernising the National Vulnerability Database for an era of AI-generated vulnerability discovery and machine-consumable security data. The RFI asks, among other things, which vulnerability-management tasks are suitable for AI, which should require human review, what safeguards belong around them, and how AI-driven decisions can remain transparent and auditable. Comments close on 13 October and are posted publicly without redaction under docket NIST-2026-0100.
An accompanying NIST blog post disclosed a tool the agency has been building called V-etalon, described as leveraging AI to aid in enriching vulnerability information. It has not been released, and the disclosure does not say how it would enter the production workflow or how analysts would review its output. Today's ITL AI webinar — “The Development of an AI Agent Enrichment Workflow at the National Vulnerability Database” — is the first public account of the approach and architecture, the problems found during implementation, and early results.
The pressure is real and nobody disputes it. CVE submissions rose 263% between 2020 and 2025; the first quarter of 2026 ran nearly a third ahead of the same quarter a year earlier. NIST enriched close to 42,000 CVEs in 2025, 45% more than in any previous year, and it was still not enough. On 15 April 2026 the NVD formally moved to risk-based enrichment, concentrating analysts on CVEs in CISA's Known Exploited Vulnerabilities catalogue, federal government software, and critical software — and reclassified roughly 29,000 backlogged records published before 1 March 2026 as “Not Scheduled”, with no committed date.
So the choice is not between an agent and an analyst. It is between an agent and a blank field. That framing is correct, it is why this is happening, and it is also exactly the framing that hides the problem.
Three fields, three failure shapes
NVD enrichment attaches three things to a published CVE: a CVSS severity score, a CWE weakness classification, and CPE applicability statements saying which product versions the flaw applies to. They are routinely discussed as one workstream because one analyst produces all three. They are not one workstream, because they are consumed differently.
CVSS fails loudly
A severity score is a number a human reads, next to a description that same human can also read. If an agent scores a trivial information disclosure at 9.8, someone in a triage meeting will say so, because the score and the evidence for it arrive in the same view. The failure mode is mis-prioritisation, which is expensive, and it is self-announcing: people notice when the queue fills with critical findings that are not critical. CVSS is also already contested — vendors, CNAs and the NVD disagree about scores routinely — so consumers have long-standing habits for treating it as an opinion rather than a fact.
CWE fails slowly
A weakness classification feeds analytics: which classes of bug appear in which parts of your estate, where secure-development effort should go, what your trend line looks like. A wrong CWE does not break any single decision. It degrades a corpus, and the damage shows up a year later as a conclusion drawn from a mislabelled population. Unpleasant, recoverable, and unlikely to be the thing that gets you.
CPE fails silently
A CPE applicability statement is not an assessment. It is a key. Your scanner takes the CPE strings it derived from your inventory, intersects them with the CPE strings on the CVE record, and emits a finding for each match. The operation is a join, and the failure mode of a join is not a wrong answer — it is an empty result.
Consider what each error does. A CPE that is too broad produces a false positive: an alert fires on something you do not run, a human investigates, and the error is discovered within a day because somebody had to look at it. A CPE that is too narrow — the wrong vendor string, a version range that stops one release short, a product renamed between major versions — produces nothing. No row, no finding, no alert, no ticket, and no artefact anywhere in your pipeline recording that a question was asked and answered “no”. The vulnerability remains in your estate and your coverage report says you are clean, which is worse than saying nothing, because it is a positive claim.
This asymmetry is not new and it is not caused by AI. Hand-written CPE has always been the flakiest part of the record, and the false-negative problem is well documented. What is new is a proposal to increase the volume of CPE production by a large factor using a process whose errors correlate — a model that misreads a vendor's naming convention misreads it the same way across every advisory from that vendor, which turns a scatter of independent mistakes into a systematic blind spot with a shape.
Why “the model is 94% accurate” is the wrong number
Any evaluation of an enrichment agent will produce a headline accuracy figure, and that figure will be an average over the three fields. Average it and you have combined a metric whose errors are caught by the next human who looks, a metric whose errors take a year to matter, and a metric whose errors are invisible by construction. The aggregate is not a summary of those three things; it is a number that conceals which of them moved.
- CPE needs recall reported separately, and against what. Precision on CPE — of the products the agent named, how many were right — is the easy direction and the one a demo shows. Recall — of the products actually affected, how many did the agent name — is the number that predicts silent misses, and measuring it requires ground truth that by definition does not exist for the records nobody enriched.
- Version ranges are not strings, and string accuracy scores them as if they were. “Correct except the upper bound” is a full miss for every asset above that bound and a full hit for every asset below it. The unit of evaluation has to be assets that would have matched, not records that look right.
- Correlated error breaks the sampling. Reviewing a random 2% of agent output estimates the error rate honestly and tells you almost nothing about whether the misses cluster on one vendor, one ecosystem, or one naming pattern. Cluster them by vendor before you report anything.
- An abstention is a result. An agent that declines to emit CPE for records it cannot resolve is more useful than one that emits its best guess, because the empty field preserves the signal that a human or another source is needed. Whether the workflow can abstain, and whether abstention is rewarded in its evaluation, is a design question worth more than a percentage point of accuracy.
The strongest objection, and why it does not close the argument
The objection is simple: 29,000 records currently have no CPE at all. An agent that gets CPE right 90% of the time is strictly better than a field that is right 0% of the time. Refusing the agent is choosing the worse option out of squeamishness about provenance.
That would be right if the alternative to a wrong value were an equally silent absence. It is not, because consumers behave differently in the two cases. A missing CPE is legible: container scanners, SCA tools and enterprise programmes already detect it and fall back — to vendor advisories, to ecosystem-native advisory databases, to package-manager metadata, to commercial enrichment feeds, to reachability analysis. Several vendors rebuilt exactly this fallback after April, publicly, and told their customers to stop treating the NVD as complete. The blank field is what triggers that behaviour.
A present-but-wrong CPE removes the trigger. The cost of the error is not one bad record; it is the suppression of the compensating control that the bad record's absence would have invoked. That is why “90% is better than nothing” is a category error: you are not comparing 90% correct against 0% correct, you are comparing 90% correct plus 10% invisible against 0% correct plus 100% visible — and the second column is the one your programme is built on.
The fix is not to refuse the agent. It is to make the two cases distinguishable again, which takes one field.
What to ask for before 13 October
The RFI asks how AI-driven decisions can remain transparent and auditable. That is the right question and it has a concrete, cheap, testable answer that belongs in the data model rather than in a policy document.
- Per-value provenance in the record, not per-record. Each enriched value — each CVSS vector, each CWE, each CPE match — carries a source: analyst, CNA, tool, or tool-reviewed-by-analyst. Per-record provenance is not enough, because the interesting records are the mixed ones where an analyst confirmed the score and a tool guessed the applicability.
- A confidence on tool-supplied values, exposed through the API as a filter. Not for humans to read. For a scanner to be configured against: “treat tool-supplied CPE below threshold X as absent” is a one-line policy that restores the fallback behaviour, and it is impossible to write today.
- “Not Scheduled” must not be silently overwritten. If a tool fills a record that human analysts were never going to reach, the record should say so permanently. A state that quietly becomes “enriched” destroys the only inventory anyone has of what was never checked.
- Publish the evaluation per field, with CPE recall broken out by vendor. Including the ground-truth construction. An aggregate accuracy figure over CVSS, CWE and CPE should be treated as a non-answer, however high it is.
- Say what the agent reads. An enrichment agent's inputs are vendor advisories, changelogs and issue trackers — text written by parties with an interest in how their product's applicability is recorded, and reachable by anyone who can file a bug. That is an untrusted-input pipeline feeding an authoritative record, and the threat model deserves a paragraph.
Comments are public and unredacted, so a short filing that says one useful thing is worth more than a long one. The single highest-value sentence available to most organisations is a concrete description of what your scanner does when cpeMatch is empty versus when it is populated — because that is the evidence that the two states are not interchangeable, and it is evidence only operators have.
What to do in your own programme this week
None of this requires waiting for NIST, and most of it is overdue independently of whether an agent ever ships.
- Make absence a distinct state in your own pipeline. If your vulnerability data model cannot express “this CVE has no applicability data” separately from “this CVE does not apply to us”, that is a schema change worth making now. It is the same change that lets you consume a provenance field later.
- Snapshot and diff. Keep your own history of NVD records. When a record acquires a CPE list it did not have last month, that is an event worth seeing — today because it means the record was enriched late, and shortly because it may mean it was enriched by a tool.
- Stop reporting coverage against the NVD alone. It stopped being a complete enrichment source in April, by announcement. A coverage metric computed against an incomplete denominator is the same failure this whole post is about, one layer up.
- If you build agents that consume vulnerability data, note what just changed. A remediation or triage agent reading the NVD has always been reading a human judgement. It is about to be reading, in part, another model's output, without a marker distinguishing the two. That is a supply-chain property of your agent, not a data-quality footnote — and it is the second time this year the industry has quietly stacked a model on top of another model's unattributed output.
FAQ
Is NIST replacing analysts with an AI agent?
Nothing announced says that. NIST has disclosed a tool under development and presented an agent enrichment workflow, and the RFI explicitly asks which tasks should require human review. The production role of the tool has not been stated, which is one of the things worth asking about before 13 October.
What is CPE and why does it matter more than CVSS here?
Common Platform Enumeration is a structured naming scheme for products and versions. NVD applicability statements list the CPEs a vulnerability affects, and scanners match those against the CPEs derived from your inventory. CVSS is an opinion a human can check; CPE is the key that decides whether the finding exists at all in your tooling.
Did the NVD stop enriching CVEs?
Not entirely. Since 15 April 2026 it enriches on a risk basis — CISA KEV entries, federal government software and critical software — while roughly 29,000 older backlogged records were reclassified “Not Scheduled”. Most CVEs are still published; fewer receive CVSS, CWE and CPE.
Would a confidence field actually help, or is it decoration?
It helps only if consumers can filter on it, which is why the ask is for API exposure rather than a note in the record. With a filter, “treat low-confidence tool CPE as missing” restores the fallback path a blank field triggers today. Without one, it is decoration.
Is agent-generated enrichment a bad idea?
No. The volume argument is sound and the backlog is real harm happening now. The claim here is narrower: the three enrichment fields should not be shipped under one accuracy number or one provenance policy, because one of them fails without producing any evidence that it failed.
How would I know if this has already hurt me?
You mostly would not, which is the point. The closest available check is retrospective: take the last ten vulnerabilities you learned about from a vendor advisory rather than from your scanner, and ask whether the NVD record had applicability data for the affected version at the time. That ratio is your silent-miss rate, measured the only way it can be.
Further reading
On this wiki:
- Vulnerability management for agent platforms — the consuming side of this data, as an engineering property.
- Vulnerability remediation agents — what happens downstream when the applicability data is wrong.
- Uncertainty & calibration — why an abstention is a result, and when to pay for one.
- Disclosure & content provenance — marking machine-generated values so downstream systems can act on the mark.
- Decision receipts and audit — the general form of “record what produced this value”.
- Discovery is not an inventory — the same shape of error, on a different asset.
Sources:
- Federal Register — RFI on Modernizing the National Vulnerability Database in the Age of Artificial Intelligence
- NIST — Shaping the NVD for the Future: We Need Your Feedback on AI-Enabled Vulnerability Management
- NIST ITL AI Webinar — The Development of an AI Agent Enrichment Workflow at the National Vulnerability Database
- NIST — Updates to NVD Operations to Address Record CVE Growth
- NVD — Vulnerability Metrics