Air-Gapped Agent Deployments

8 min read

O31
Operation · AgentOps: Deploy & Operate

Air-gapped agent deployments.

The question everyone asks about an air gap is where the model will run, and that is the one question with a vendor answer. The expensive surprises are the half-dozen internet dependencies your agent had without anyone deciding it should: the package index it reads version numbers from, the tool registry it discovers servers through, the hosted judge your eval suite calls, the error reporter that silently drops, the certificate revocation list that never resolves. Inventory those before you move, because inside the gap each one fails as an unexplained quality regression rather than as a connection error — and plan on a different model tier, which makes this an eval migration wearing an infrastructure migration's clothes.

STEP 1

Name the three claims separately, because they cost different amounts.

"Air-gapped" gets used for three postures with different engineering consequences, and teams routinely buy one while promising another. Separate them in writing before the architecture review, because the control that satisfies a residency clause does not satisfy an isolation requirement.

# Three postures, in increasing order of what breaks

SOVEREIGN CLOUD   provider infrastructure, contractual region + jurisdiction
                  internet present; dependencies all still work
                  buys: residency, legal venue        breaks: nothing technical

SELF-HOSTED       your infrastructure, egress allowed through a policy
                  buys: custody of weights + data     breaks: provider features

AIR-GAPPED        no route to the internet at all, by design
                  buys: isolation you can demonstrate
                  breaks: every implicit network dependency, all at once

# The sentence worth writing down before you commit

  "we need X" where X is one of the three, not the word air-gapped

Only the third posture is the subject of this page. The first two are addressed by data residency and sovereignty and self-hosted inference for agents, and if a residency requirement is what you actually have, stopping at posture one is a legitimate and far cheaper answer.

STEP 2

The model tier available inside is the binding design constraint.

This is the part that changes the product rather than the plumbing, and the vendor matrices make it explicit rather than hiding it. IBM's self-hosted Bob, generally available on 1 October 2026 for on-premises, private-cloud, sovereign-cloud and air-gapped environments, keeps the harness intact — the shell, parallel tool calling, skills, operating modes — while the models supported for customer-managed infrastructure at general availability are NVIDIA Nemotron and Poolside Laguna. The hosted and hybrid configurations are where Claude Sonnet 5.0 and Opus 4.8, Gemini 3.7 Flash and GPT 5.6 Sol appear. The harness crosses the gap; the model does not.

Three operational consequences follow, and all three are schedule items.

  • Your prompts and skills are untested assets. Everything calibrated against a hosted frontier model has to be re-earned against the model you can run inside — this is exactly the problem in prompt portability, arriving as a deployment decision rather than a model choice.
  • Re-scope the task set, not just the prompts. If the on-premises model needs more steps for the same task, your step ceilings, timeouts and context budgets were all sized for a different model. A longer loop at the same ceiling reads as a mysterious rise in truncated runs.
  • Capacity is now yours to plan. No burst, no elastic quota, no provider absorbing your peak. Size for p99 concurrency and accept idle GPUs, or implement admission control and accept queueing — see serving agent traffic.

Run the comparison before the migration, not after. Stand the candidate on-premises model up in your normal environment, point your existing task suite at it, and read the gap while you still have both. Teams that discover the quality delta after the gap is closed cannot tell a model regression apart from a broken dependency, because inside the gap both present as "it got worse".

STEP 3

Enumerate the silent dependencies; most of them are not the model.

An agent is a program that looks things up. Cut the network and the lookups do not raise exceptions — the model answers from memory instead, which is the worst available failure mode because it is indistinguishable from working. Walk this list against your own agent and mark each row with what replaces it.

# Implicit network dependencies, and how each one fails inside the gap

package index / registry    version lookups -> model guesses a version
MCP / tool registry         server discovery -> stale local catalog only
docs + web retrieval        grounding -> confident answers from weights
hosted LLM judge            eval scores -> suite silently skips or errors
telemetry + error reporting drops -> you lose the one thing you need most
model + container updates   no pull -> a release process, see Step 5
CRL / OCSP / NTP            TLS + token validation fails, intermittently
license activation          phones home -> expiry outage on a quiet Sunday

Two rows deserve emphasis. The grounding row is the one that produces wrong work rather than failed work: an agent that could previously read current documentation now answers from a knowledge cutoff and has no way to signal the difference, so internalise retrieval deliberately — a mirrored docs corpus, an internal package index, a tool catalogue you refresh on a schedule — and make the agent's retrieval tool fail loudly when the corpus is older than a declared freshness bound. And the time row is the one that produces a 3am incident: certificate validation and token expiry both depend on clocks and revocation data that an isolated network has to supply itself.

Test the inventory the hard way. Block egress in a staging environment that is otherwise identical and run your full task suite — the same technique as fault injection for agents. Every dependency you did not know about shows up in one afternoon, which is considerably cheaper than finding it during a sovereignty audit.

STEP 4

Evaluation and observability have to move inside, not be skipped.

The predictable casualty of an air gap is the eval suite, because it is the one piece of the stack whose dependencies are mostly hosted: a judge model at a provider, a managed tracing backend, a dataset in object storage somewhere else. The pattern to avoid is the common one — the suite gets disabled "temporarily" for the air-gapped deployment, and the environment with the least observability becomes the one running with no quality gate at all.

  • Host the judge inside. A smaller open-weight judge running locally, re-calibrated against your existing labelled set so you know the agreement rate before you trust it. The calibration step is not optional; a different judge is a different metric. See LLM-as-judge for agents.
  • Prefer rule-based graders where they are possible at all. Inside a gap, every graded assertion you can express as an exact check instead of a model judgement removes a dependency and a calibration argument at once.
  • Traces stay and therefore accumulate. Nothing ships them off-site, so retention is a disk-capacity decision and a disclosure decision in the same breath — set it explicitly, as in trace sampling and retention.
  • Decide how findings get out. Someone will need to act on an aggregate metric from outside the gap. Define that export as a reviewed artifact — a signed summary, no raw traces, on a named cadence — rather than leaving it to whoever has a USB stick.

And run the eval suite inside the gap on the real deployment, not only in the staging environment with egress. The configuration differences between those two are precisely where the air-gap-specific failures live.

STEP 5

Updates become a release process with a human in the loop.

Outside a gap, patching is a background activity. Inside, every model version, container image, tool server, dependency and CVE feed arrives through a deliberate, authenticated import — and the default outcome of a deliberate process nobody scheduled is that nothing gets updated for eleven months.

  • Put a cadence on the import, in writing, with an owner. Monthly is a reasonable default for images and dependencies; model versions can be slower, but "when someone asks" is not a cadence.
  • Verify everything on the way in by digest and signature. The transfer medium is the supply chain now, and the import step is the single highest-value place to enforce pinning and verification. An unsigned image that crossed the gap on a laptop has had a trust decision made for it by logistics.
  • Mirror the vulnerability feeds too. Isolation does not patch anything; it only removes your notification channel. An air-gapped agent platform that cannot see advisories is running known-vulnerable components with no way to know — the inventory half of vulnerability management for agent platforms.
  • Treat a model update as a release, with the eval gate from Step 4. This is the one upside of the gap: nothing changes under you. No silent endpoint upgrade, no unpinned vendor default moving on a Tuesday. You own the version, which also means nobody else will notice the regression for you.

Write down the exception path before you need it, because you will need it. When a production incident needs a vendor's help and the vendor cannot see anything, the improvised answer is someone screenshotting logs to a phone. Decide in advance what may leave, redacted how, approved by whom — and note that the same reasoning applies to the agent's own traces, which contain source code and customer data in the environments where people choose air gaps.

STEP 6

The acceptance gate, and the number that tells you whether it worked.

An air-gapped deployment is accepted on evidence, not on a feature checklist, because the feature checklist is the thing that ported cleanly. Five checks, each a yes or a no.

  • The task suite passes inside the gap, at a pass rate you wrote down in advance as acceptable with the on-premises model — not at parity with the hosted deployment, which you will not get and should not promise.
  • Egress is denied by default and the denial is logged, so an attempted lookup is a visible event rather than a missing one. This is the inside-out version of egress control for agents: the policy is trivial, the logging is the product.
  • Every retrieval corpus declares its freshness and the agent refuses, or flags, when it is stale past the bound.
  • The eval suite and the judge run inside, with a recorded agreement rate against the labelled set you had before.
  • An import has actually been performed end to end — one full cycle of model, image and feed update, timed, with the signature checks exercised. An import path that has never run is a plan.

Then watch one number in production: the share of runs that complete without any tool call returning a stale-corpus or unavailable-dependency signal. It moves first and it moves early, well before quality metrics degrade visibly, because an agent working around a missing lookup still produces an answer. Where this posture is being adopted for governance reasons, the same evidence feeds the agent inventory entry and any impact assessment that follows.

Before committing to an air gap, run the cheap experiment: in a staging environment identical to production, block all egress and run your full task suite for one day. Count two things — how many tool calls failed, and how many runs produced a confidently wrong answer instead of failing. The second count is the real cost of the air gap, it is almost always larger than teams expect, and it is the only number that tells you how much retrieval you have to internalise before the gap is a safe place to operate rather than a compliant one.