AI Blog

Same weights, different refusals: Argon ships its guardrails as an entitlement

Google released Gemini 4 Argon to vetted Fairwind defenders with the cyber guardrails switched off, enforced by org verification, phishing-resistant MFA, team-scoped access and per-employee usage records. That is the first version of capability gating that could actually hold — and it means a model identifier no longer names a behaviour.

By Agentic AI Wiki 15 min read

Two organisations can now call the same Google model and get materially different answers to the same question, and nothing in the API tells either of them which version they are holding. Gemini 4 Argon went out on 30 September 2026 to vetted members of Google's Fairwind programme with the cyber guardrails switched off — not loosened by prompt, switched off as a property of the entitlement, granted after Google verifies the organisation and only in exchange for phishing-resistant MFA, access confined to security, incident-response or pen-test staff, and a per-employee record of who used it. That is the first implementation of capability gating on this axis that could actually hold, because an entitlement cannot be talked around the way a refusal can. It also quietly retires an assumption every agent stack is built on: that the model string identifies the behaviour.

At a glance

The release is unusual in three separate ways, and only one of them is about the model.

DimensionWhat Google shippedWhy it is unusual
Who gets itFairwind-vetted defenders and Google's own security teamsFrontier release gated by organisational vetting, not by tier or spend
What they getThe same model with cyber guardrails offThe staged artefact is the policy, not the weights
ConditionsOrg verification, phishing-resistant MFA, team-scoped access, usage recordsEnforced outside the model, where it cannot be rephrased away
Everyone elseBroader rollout "after additional testing", no dateA flagship nobody can buy, announced anyway
Limits2M-token context, up to 1M output tokensOutput cap up from roughly 64K — a 15× change
PriceIntroductory $2 / $10 per million in / outAnnounced to move to $4 / $20 after the introductory period

Fairwind itself launched on 2 September 2026 with more than 650 partners — governments and national cyber authorities, critical-infrastructure operators, technology platforms and security vendors — initially carrying Gemini 3.8 Flash Cyber and CodeMender, Google's vulnerability-discovery-and-patch system. Argon is the upgrade to that channel, and Google says it reaches frontier performance on software engineering, legal and financial knowledge work, and cyber defence.

Gating moved from the refusal to the entitlement

One set of weights served through two policy paths The weights are the same on both paths. The gate is outside the model. ONE CHECKPOINT gemini-4-argon weights PATH A — ANY PAYING CALLER + cyber refusal layer rephrasable, version-drifting, unauditable from outside PATH B — FAIRWIND ONLY no cyber refusal layer same weights, different served policy FOUR GATES, NONE OF THEM PROMPTABLE 1 — verified organisation, security and ethics record checked 2 — phishing-resistant MFA on every account that touches it 3 — access limited to security, IR or pen-test employees 4 — per-employee record of who used it, retained WHAT AN ATTACKER CAN REACH a brilliant jailbreak gets you Path A, degraded it does not get you a Fairwind grant An entitlement cannot be talked around. That is the whole improvement — and the reason the model string stops naming a behaviour.
The weights are identical on both paths. What differs is a policy selected by an entitlement nobody can prompt their way into.

The standard criticism of safety refusals is that they are the weakest control in the stack: a refusal is a policy applied at generation time over a capability that is still present, so it degrades under rephrasing, it drifts between versions without a version number changing, and it cannot be audited. Anyone who has tried to hold a real boundary with one has learned to put the control somewhere the model cannot argue with it.

Fairwind is that advice taken seriously at the distribution layer. The four conditions Google attaches — a verified organisation with a checked security and ethics record, phishing-resistant MFA on accounts that touch the model, access limited to employees working directly in security, incident response or penetration testing, and per-employee usage tracking — are all enforceable without the model's cooperation. An attacker who writes a brilliant jailbreak does not acquire a Fairwind grant. The control is an access-control problem, which is a solved genre, rather than a model-behaviour problem, which is not.

It is also not a one-off. OpenAI reached the same structure a month earlier from a different direction: having classified its new model at the Critical cybersecurity threshold on 1 September, it routed first access through Daybreak, an application-based cyber programme, rather than the normal tiers. Two labs, independently, arrived at the same answer — that the right place to gate a dual-use capability is the population of callers, not the sampling loop.

Read this as a genuine improvement, because it is one. For years the public debate over dual-use model capability has been argued as if the only dial were how hard the model says no, which is a dial with a known ceiling. Moving the decision to a vetted population with contractual conditions and an audit obligation is what every other dual-use industry already does, and it is the first version of this that a security engineer can reason about.

The model string stopped naming a behaviour

What a model identifier used to pin down, and what it pins down now Three columns. Before entitlement-gated policy tiers, a model identifier was treated as pinning the weights, the refusal policy and therefore the behaviour, so eval results, questionnaire answers and incident reports could all be keyed to it. Now it pins the weights only: two callers passing the same string can be served different refusal policies. The third column gives the replacement key — model identifier plus policy tier plus system prompt version plus tool list — which is the smallest set that a result can honestly be filed under. The behavioural primary key just lost a column WHAT WE ASSUMED The ID pinned behaviour weights: yes refusal policy: yes so: two callers, same string, same answers evals, questionnaires and incidents all keyed to it WHAT IS TRUE NOW The ID pins the weights weights: yes refusal policy: no entitlement selects the served policy an unreproducible eval now looks like harness drift THE REPLACEMENT KEY Four fields, not one model identifier policy tier system prompt version tool list anything you cannot rebuild from these was never pinned One extra column, written once, converts a class of mysterious results into explicable ones.
The identifier still pins the weights. It no longer pins the policy, and the policy is what your users experience.

Here is the cost, and it is paid by everyone, including the organisations with no interest in cyber capability. Every agent stack treats the model identifier as the behavioural primary key. Your eval results are filed under it. Your vendor security questionnaire answers it. Your incident write-up names it. Your reproducibility claim depends on it. All of that silently assumed that two callers passing the same string get the same policy — and for this model, on this release, they demonstrably do not.

The practical failures are mundane and specific. An eval suite run by a Fairwind-entitled security team produces refusal-sensitive numbers that a sister team outside the entitlement cannot reproduce, and the discrepancy looks like harness drift. A questionnaire answer that says "we use Gemini 4 Argon" no longer conveys what the model will decline, which was the only reason the question was asked. A vendor comparison built on published benchmark scores is comparing an entitlement tier it cannot buy — the scoreboard numbers come from somewhere, and "somewhere" is now a parameter.

The fix is cheap if you do it before you need it: record the policy tier beside the model identifier everywhere you record the model identifier. Treat it as part of the configuration that a result is keyed to, in exactly the way you already treat the system prompt version and the tool list — see reproducibility and determinism for why the key has to include everything that moves. One extra column, written once, makes a class of unreproducible results explicable instead of mysterious.

What Google did not publish matters more than the benchmark it won

Coverage of Argon has centred on scores. Third-party indices put it at or near the top of general-capability leaderboards, and it reportedly ties for first on a vulnerability-remediation benchmark. Set all of that aside, because the operationally important sentence in the announcement is one that is not there.

Google does not say which restrictions come off for Fairwind members, and it does not say what the guardrails stop the model doing while they are on. The Frontier Safety Framework names cybersecurity as one of four risk domains and describes the intent — refuse harmful cyber and CBRN requests while preserving legitimate dual-use research — but the delta between the two served policies is not characterised anywhere a customer can read.

That absence has a measurable consequence. Without a published diff, nobody outside Google can tell whether a given failure is a capability limit or a policy limit. Those two have opposite remedies: a capability limit means change the model or the scaffold, a policy limit means change the request, the entitlement, or the vendor. Teams that cannot distinguish them will spend weeks on prompt engineering against a refusal, which is the single most common waste in this category of work — and it is why the enumerated refusal reason is a product feature and not a nicety.

The honest counter-argument: publishing the diff is publishing a map of what the de-guardrailed tier unlocks, which is uncomfortably close to a capability advertisement. That is a real tension and there is no clean resolution. But there is a middle path nobody has taken — publish the categories and their refusal rates on a fixed public probe set for both tiers, without publishing which prompts pass. Customers get a measurable delta; attackers get a histogram.

The two numbers that change your harness rather than your scoreboard

Output token cap and the cost of one cap-filling response A horizontal bar chart on a logarithmic footing comparing output ceilings. Earlier Gemini models capped output at roughly 64,000 tokens, so one cap-filling response cost about 64 cents at ten dollars per million output tokens. Gemini 4 Argon raises the ceiling to one million tokens, so one cap-filling response costs about ten dollars at the introductory rate and about twenty dollars at the announced post-introductory rate of twenty dollars per million. The worst single call in a system therefore moves by roughly fifteen times, in a dimension most per-task budgets do not measure. One response that fills the cap OUTPUT CEILING COST OF ONE FULL RESPONSE Earlier Gemini 64K cap, $10/M out 64,000 $0.64 Argon, introductory 1M cap, $10/M out 1,000,000 $10.00 Argon, post-introductory 1M cap, $20/M out 1,000,000 $20.00 The ceiling was doing unacknowledged work: it was the thing that terminated a runaway generation. A 15× cap turns a truncation error into an open-ended run, and most per-task budgets count calls or input tokens — neither of which changed. Bars are not to scale against each other beyond the 64K-versus-1M contrast; the cost figures are the point.
A single response can now cost more than a thousand ordinary ones. The cap is a budget decision, not a length decision.

Two of the published figures do real work on an agent stack. The first is the output limit: up to one million output tokens, against roughly 64,000 on earlier Gemini models. That is not a convenience. Every timeout, retry policy, streaming buffer, cost guard and per-task budget in your harness was calibrated against an output ceiling three orders of magnitude lower, and the thing that used to bound a runaway generation was the cap itself. Raise the cap and a clean truncation error becomes an unbounded run — the failure mode changes from "malformed JSON" to "the task is still going and the bill is open".

The arithmetic is worth doing once. At the introductory rate of $10 per million output tokens, one response that fills the cap costs $10. At the announced post-introductory rate of $20, it costs $20. A 64K response at the same rates costs $0.64 and $1.28. So the new ceiling moves the worst single call in your system by a factor of roughly fifteen, in a dimension your per-task cap probably does not even measure — most caps count calls or input tokens, both of which are unchanged. If you take one action from this post, make it a per-task output-token budget enforced in the harness, as in cost control in the loop.

The second number is the price path. Introductory $2 and $10 per million input and output tokens, announced to double to $4 and $20 when the introductory period ends, with cached input at a steep discount. A price with an expiry is a term, not a rate, and unit-economics models built on it carry a 2× step function that nobody wrote down. That is the same discipline as writing the price path down: record the rate, the date it was quoted, and the date it changes, in the model itself.

What to do, on either side of the vetting

Split the advice, because the two populations have nothing in common.

  • If you have or want a Fairwind grant: the conditions are the work, not the application. Team-scoped access and a per-employee usage record mean you need a named population, an authentication posture you can attest to, and a retained log that survives an audit — read them as the requirements they are, alongside access reviews for agent credentials. Assume the grant is revocable on exactly these conditions.
  • If you will never have one: your exposure is the mirror image. Your refusal profile can also change without a version bump, in the restrictive direction, and almost nobody measures it. A refusal inside an agent loop is not a visible "no" — it is a tool call that does not happen and a plan step that quietly disappears. Start measuring it, per the recipe in refusal monitoring in production.
  • Everyone: add the policy tier to the key. Model ID, policy tier, system prompt version, tool list. Anything you cannot reproduce from those four was never reproducible.
  • Everyone: cap output tokens per task in the harness, not in hope. The provider's ceiling is no longer a safety net.
  • Everyone: stop treating a published benchmark score as a property of a name. Ask which tier produced it, and whether you can buy that tier. Increasingly the answer to the second question is no — which is the real content of reading benchmarks in 2026.

And one prediction worth holding loosely: if entitlement-gated policy tiers work — if Fairwind and Daybreak produce defensive value without obvious abuse — this becomes the default shape for every frontier capability that is dual-use, which is most of the interesting ones. In that world the question "which model are you using" is permanently insufficient, and the vendor questionnaires, the eval harnesses and the incident templates that still ask it will all need a second field.

FAQ

Can my company get Gemini 4 Argon?

Only through Fairwind today, and only if you are a government or national cyber authority, a critical-infrastructure operator, a core technology platform or a security provider that passes Google's vetting. Google says a broader rollout follows additional testing, with paid API customers and top-tier subscribers first, and has not given a date.

Is "guardrails off" the same as an uncensored model?

No. What Google describes coming off is the cyber-specific refusal layer for vetted defensive work. The Frontier Safety Framework covers four risk domains, and nothing in the announcement suggests the other three are affected. Google has not published the exact delta, which is the gap this post is about.

Does the entitlement model actually stop misuse?

It changes the attack from a prompting problem into an impersonation-or-insider problem, which is strictly harder and, unlike a jailbreak, leaves an audit trail with a name on it. It does not stop a compromised Fairwind partner, which is why the MFA and usage-record conditions are the load-bearing parts.

Why does the output cap matter more than the context window?

Because the context window bounds what you send and your code already measures that, while the output cap bounds what the model generates and most harnesses relied on the provider's low ceiling to terminate runaway generations. A 15× cap increase removes a limiter that was doing unacknowledged work.

What is the one thing to change this week?

Add a policy-tier field next to every recorded model identifier — in eval metadata, trace attributes and questionnaire answers. It costs an afternoon and it is the prerequisite for every other conclusion here.

Further reading

On this wiki:

Sources: