The United States now has a pre-release review gate for frontier AI models, and on 4 August the White House told the labs sitting in the room that the framework behind it will not be published. Strip away the politics and what is left is an engineering object this field already knows how to judge: a benchmark. Except this one has no published methodology, no disclosed threshold, no reported results and no route to contest a score — which removes every check that makes a benchmark number mean anything. A benchmark you cannot read is not a weaker benchmark. It is an assertion.
What actually happened
Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security, was signed on 2 June 2026. It does two things at once, and the second is the one that matters here.
| Element | What the order says | Status |
|---|---|---|
| Classified benchmarking process | Treasury, the Secretary of War (through the NSA) and Homeland Security (through CISA) build a classified process to assess models' advanced cyber capabilities. | Due within 60 days — by 1 August 2026. |
| The threshold | That same process sets the capability level at which a model is designated a covered frontier model. | Unpublished. |
| Pre-release access | Developers are invited to engage voluntarily and, if their model is covered, to give the government access for up to 30 days before release, under confidentiality, cybersecurity, insider-risk and IP protections. | Voluntary. |
| Scope | Reporting on the framework describes a covered model as closed-weight and state-of-the-art. Open-weight releases sit outside it. | Excluded. |
The 1 August deadline passed without a Federal Register notice, a NIST or CISA publication, or a statement from the Office of Science and Technology Policy. Three days later the administration convened Meta, Nvidia, Microsoft, OpenAI, Anthropic and a set of smaller companies to walk them through the framework, and confirmed it would stay unpublished — visible only to the companies that choose to participate.
None of this is a surprise on the substance. The frontier labs and the government's evaluation body at NIST have been sharing pre-release model access informally for a while; what changed is that an informal courtesy acquired a name, a threshold and a designation.
The gate, and where the visibility line falls
Every stage before release is invisible
A developer approaches the government, the model is measured against a classified process, a threshold nobody outside has seen decides whether it is covered, and if it is, agencies get up to a month with it. The public learns none of this. There is no requirement to disclose that a review occurred, what it measured, what the model scored, or whether the threshold was crossed. The first externally observable event in the entire sequence is the model going on sale.
The secrecy has a real argument behind it
It is worth stating fairly, because it is not obviously wrong. A published benchmark for offensive cyber capability is a published curriculum for building one, and a disclosed threshold tells an adversary exactly how close a model has to be. Both concerns are legitimate. The question is not whether some of this should be classified — clearly some of it should — but whether all of it needs to be, and the answer to that determines whether anyone outside the room can reason about the result at all.
What a benchmark needs to mean anything
This wiki spends a lot of pages on how to read an eval, and none of that advice is about trusting the people who ran it. It is about the checks a reader can perform. Line them up against the three kinds of evaluation and the difference is not one of degree.
Contamination. A public benchmark's biggest failure mode is that its items leaked into training data, and the way you detect that is by inspecting the items. An internal eval controls contamination by construction, because you wrote it from your own traffic. A classified benchmark can be contaminated exactly like any other — the frontier labs train on the same open internet the test authors drew from — and no participant can check. The government would have to detect its own contamination, and would have no reason to announce it.
Variance. An agentic score is a sample, not a measurement. Run the same harness twice against the same weights and you get different numbers; the spread on multi-step tasks is wide enough that a single run routinely cannot distinguish two models. Published methodology tells a reader how many runs, at what temperature, with what pass criterion. Absent that, "the model scored above the threshold" is a sentence with no error bar, and a threshold applied to a noisy statistic is a coin weighted by whoever picked the sample size.
Reproduction. The reason anyone believes a benchmark result is that somebody else re-ran it and got roughly the same thing. Here there is exactly one party who can run the test, and the same party sets the threshold and makes the designation. That is not a criticism of their competence; it is a structural observation about a measurement with no second observer.
Contestability. A lab told its model is covered has no published criteria to argue against. It cannot say "your task set is contaminated", "your variance is larger than your margin" or "your scaffold is unrepresentative", because it has not seen any of those things. The process is voluntary, so the formal remedy is to withdraw — which is not a remedy at all for a company that sells to the federal government.
The honest summary: participating labs are being asked to accept a capability designation produced by a test they cannot inspect, scored by a method they cannot reproduce, against a threshold they cannot see. Every one of those would be a disqualifying objection in a model card. Here they are the design.
The unit problem: capability is not a property of the weights
Set the secrecy aside for a moment, because there is a second problem that would remain even if the whole framework were published tomorrow. The thing being gated is a set of weights. The thing being worried about is what an agent built on those weights can do. Those are not the same object, and the gap between them is not small.
Measured agentic capability is a function of the scaffold as much as the model. The same weights, given a better tool set, a longer iteration budget, a working memory across attempts and a harness that retries intelligently, solve tasks they failed on at release. This is the well-documented gap between what a model can do and what a particular elicitation gets it to do, and it is precisely why agent evaluation grades trajectories rather than models. A pre-release review measures one configuration of a system whose other components will be replaced by third parties within weeks.
The practical consequence is that the dangerous configuration usually does not exist on the day of the review. It is assembled later, in public, by an open-source project wiring a released model into a better loop — and there is no stage in this framework at which that assembly is observed. Gating the weights is like certifying an engine and never looking at the vehicle.
This is not an argument that the review is worthless. Some capabilities really are latent in the weights and a pre-release look is the only chance to see them before distribution. It is an argument that a model-level gate is a weak instrument for an agentic risk, and that the framework's own logic — measure cyber capability, designate, review — quietly assumes a stability the field stopped having in about 2024.
Open weights sit outside the line, and that is the load-bearing exclusion
Reporting on the framework consistently describes covered models as closed-weight. If that holds, the gate binds a distribution channel rather than a capability, and three things follow.
Equivalent capability, unequal treatment. An open-weight release at the same measured capability as a covered model goes out with no review at all. If the review is worth doing, this is a hole in it. If the review is not worth doing, it is a tax on one group of companies. Both readings are uncomfortable and the framework has to pick one.
A structural timing advantage. A closed lab that engages accepts up to thirty days of pre-release latency it cannot compress. An open-weight competitor does not. Thirty days is a meaningful lead in a market where releases land monthly, and it accrues to whoever opted out of the oversight — which is an odd incentive to build deliberately.
The threshold cannot be enforced downstream anyway. Once weights are public they are fine-tunable, and safety post-training is the cheapest layer of the stack to remove. Any capability threshold defined on a released open-weight model is a statement about the checkpoint, not about what anyone will be running a fortnight later. Excluding open weights may be less a policy choice than an admission that the instrument does not work there.
Note what this does to the phrase "voluntary". A framework is voluntary in the sense that no statute compels participation. It is not voluntary in the sense that matters to a company whose largest customer is the federal government and whose competitors are in the room. Secret criteria plus procurement leverage plus no appeal is a regulatory regime with none of the procedural protections that word normally implies.
What changes if you build agents
Nothing at your API boundary today. But three second-order effects are worth planning around.
Release timing acquires a tail you cannot see. A model you are waiting on may be sitting in a thirty-day window, and you will not be told. If your roadmap assumes a launch date announced by a vendor, add slack — and keep your migration path warm, because the mitigation for unpredictable release timing is not depending on a single provider's calendar.
Your own evals are now the only public evidence about a model. If capability assessment moves behind a classification boundary, the externally verifiable numbers are the ones you and other practitioners produce. That raises the value of a small, honest, task-specific eval set considerably, and it is a good week to reread reading agent benchmarks and to actually publish yours.
The open-weight side of the line is where the unreviewed capability accumulates. If you self-host, you are running models that no gate examined, which is a reason to own the controls yourself — sandboxing, egress policy, scoped credentials — rather than inheriting them from a provider's safety layer. That is good practice regardless; it is now also the thing standing where the review is not.
Three disclosures that would cost nothing
The strongest version of the government's position is that the test contents must stay classified. Grant it entirely. Everything below survives that constraint, and each one restores a check the current design removes.
Publish the methodology without the items. How many runs per task, what scaffold, what pass criterion, how variance is handled, how contamination is screened. This is what a model card discloses and it leaks no capability. Its absence is the difference between a measurement and a claim.
Publish the threshold's units and an aggregate count. Not the number — the shape of the number, and how many models were reviewed and how many were designated in a given period. A single integer per quarter tells the public whether this is a gate that ever closes, and tells an adversary essentially nothing.
Give designated developers a written basis and a route to contest it. Even a classified basis, delivered under clearance, is more than exists now. A capability finding with commercial consequences and no appeal is a procedural gap, not a security requirement.
The uncomfortable part is that these are the same three things this field demands of every lab that publishes a benchmark result. The standard we hold vendors to is methodology, error bars and reproducibility. It is a strange moment to accept less from the only evaluator whose verdict can hold a release.
FAQ
Is the frontier model review mandatory?
No. Executive Order 14409 establishes a voluntary framework — developers are invited to engage and to offer pre-release access for up to 30 days. There is no statutory compulsion. In practice, participation is shaped by the fact that the government is a major customer and that the largest labs were all in the room.
What is a "covered frontier model"?
A model whose advanced cyber capabilities exceed a threshold set by a classified benchmarking process built by Treasury, the NSA and CISA. The threshold itself has not been published. Reporting on the framework describes covered models as closed-weight and state-of-the-art, which places open-weight releases outside it.
Was the framework ever published?
No. The executive order's 60-day deliverable fell due on 1 August 2026 with no Federal Register notice or agency publication, and on 4 August the administration briefed AI companies on the framework while confirming it would not be released publicly.
Does this delay model releases?
Potentially by up to thirty days for a participating developer whose model is designated as covered. Nothing requires either the delay or its disclosure to be announced, so from the outside a review window and an ordinary schedule slip look identical.
Why does an unpublished benchmark matter more than an unpublished audit?
Because a benchmark produces a number that gates a decision, and a number's meaning comes entirely from its methodology. Contamination, run count, variance and pass criteria are not footnotes to a score — they are what makes a score comparable to anything. Without them the output is a designation, and a designation is only as good as trust in the designator.
Does this replace the EU AI Act for models sold in Europe?
No. They are separate regimes with different triggers and different obligations, and a model can be subject to both. The AI Act's general-purpose and high-risk duties turn on documentation, transparency and logging rather than on a classified capability threshold.
Further reading
On this wiki:
- Benchmark Contamination — why an unreadable test set is an unfalsifiable one.
- Eval Variance & Statistical Power — what a threshold applied to a noisy score actually decides.
- Reading Agent Benchmarks — the checks this framework removes.
- The Regulatory Landscape — where a voluntary US framework sits next to everything else.
- Open-Weight vs Closed Models — the line the exclusion is drawn along.
Sources:
- Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security — The White House
- New Executive Order Addressing Early Government Access to Frontier AI Models — WilmerHale
- White House to host AI companies to review new model-testing framework — CNBC
- White House won't publicly release the AI model evaluation framework — Fortune
- Trump AI framework excludes open AI models — Axios
- Five Questions the US Government Should Answer About Its Secretive Frontier AI Framework — Tech Policy Press