AI Blog

The harness crossed the air gap; the model did not

IBM made self-hosted Bob generally available on 1 October 2026 for on-premises, private-cloud, sovereign-cloud and air-gapped environments, with the shell, parallel tool calling, skills and modes intact. The models supported on customer-managed infrastructure are NVIDIA Nemotron and Poolside Laguna — not the hosted Claude, Gemini and GPT options. The feature list ports; the behaviour has to be re-earned, which makes a sovereignty migration an eval migration wearing infrastructure clothes.

By Agentic AI Wiki 12 min read

Read IBM's two lists next to each other and the sovereignty trade stops being abstract. Self-hosted Bob went generally available on 1 October 2026 for on-premises, private-cloud, sovereign-cloud and air-gapped environments, and it keeps the agent harness whole — the shell, parallel tool calling, skills, operating modes. The models supported on customer-managed infrastructure at general availability are NVIDIA Nemotron and Poolside Laguna; the hosted and hybrid configurations are where Claude Sonnet 5.0 and Opus 4.8, Gemini 3.7 Flash and GPT 5.6 Sol appear. The portable artifact is the one your migration checklist covers. The artifact that decides whether the output is any good is the one that stayed behind — which means the air-gap project you costed as infrastructure is a re-validation project, and the thing you have to rebuild is your eval suite.

At a glance

Three postures get called the same thing in procurement documents, and they have different prices.

PostureWhere the code sitsModel you can runWhat you must re-validate
Sovereign cloud Provider infrastructure, contracted region and jurisdiction Hosted frontier models, unchanged Essentially nothing technical
Self-hosted, hybrid model Your infrastructure; prompts still leave for inference Hosted frontier models via an external service Egress policy, latency budgets, failure handling
Self-hosted or air-gapped, local model Code, development context and build artifacts stay inside Customer-managed open-weight models — for Bob at GA, Nemotron and Laguna Prompts, skills, step budgets, the eval suite, the cost model
What changes across three deployment modes of the same agent platform Three columns for sovereign cloud, self-hosted and air-gapped deployment. In all three the agent harness box is identical. The model box changes from hosted frontier models to on-premises open-weight models. The external lookup channel is present in the first two columns and severed in the air-gapped column. SOVEREIGN CLOUD SELF-HOSTED AIR-GAPPED HARNESS shell, tools, skills identical shell, tools, skills identical shell, tools, skills identical MODEL hosted frontier what you built against hosted or local hybrid is a choice local open-weight a different agent LOOKUPS docs, registries work docs, registries policy-gated docs, registries severed no route out, by design The only row that is genuinely unchanged is the top one — and it is the row the product page describes. The middle row decides output quality. The bottom row decides whether a missing lookup fails loudly or becomes a guess.
One row is genuinely unchanged, and it is the row the product page describes.

The feature list ported; the behaviour did not

IBM's framing is accurate and it is also the source of the confusion. Bob is an agentic software-development platform — understanding code, planning work, executing changes, validating results — and the self-hosted edition genuinely keeps the capabilities that make it that: the shell, the IDE experience, parallel tool calling, skills, modes. If you listed what your developers interact with, almost all of it crosses the gap.

But none of those capabilities decide whether a patch is correct. The harness decides what the agent can attempt; the model decides how often the attempt lands. Swap the model underneath an identical harness and you have changed the only variable that the user experience is actually a function of, while every artifact that mediates between you and the model — system prompts, skills, tool descriptions, step ceilings, retry policies, acceptance thresholds — stays pinned to a model you are no longer running.

Those artifacts are not configuration. They are measurements. A mature agent prompt is a compressed log of one model's failure distribution: this clause exists because that model dropped file paths, this one because it over-edited, this threshold because its confidence was miscalibrated in a particular direction. Carry them across and you pay twice — tokens spent suppressing failures the new model does not have, and no clause at all for the failures it does. That is the general problem in prompt portability; an air gap is just the version of it you cannot postpone.

Read the model matrix as the price list

What crosses an air gap and what does not Three columns. The harness crosses intact, carrying the shell, parallel tool calling, skills and modes. The model does not cross: the models supported inside are open-weight ones rather than the hosted frontier options. External lookups are cut entirely and must be mirrored inside or they become guesses. CROSSES INTACT The harness shell, parallel tool calls, skills, operating modes this is the feature list on the product page DOES NOT CROSS The model inside: open-weight, customer-managed every prompt and eval calibrated elsewhere CUT ENTIRELY The lookups docs, package index, tool registry, judge mirror them inside, or they become guesses Only the first column is what a migration checklist usually covers. The second and third are where the work is, and both land on the eval suite.
What crosses, what does not, and which column your migration plan is about.

The practical reading advice for any vendor's sovereign or air-gapped offering is to skip the capability page and go to the supported-model matrix, because that table is the specification. If the models certified for customer-managed infrastructure differ from the hosted list — and at Bob's general availability they do — then the vendor has told you, precisely and in public, which model tier your isolation requirement costs you.

This is not a criticism of the arrangement; it is the only honest way to ship it. An air-gapped deployment can only run weights the customer is licensed to hold and has hardware to serve, which rules out every model whose weights do not leave the lab. NVIDIA Nemotron and Poolside Laguna are in that list because they can be, and a vendor that promised frontier-hosted quality inside a gap would be promising something nobody can deliver.

What follows is a reframing of the project plan rather than a reason to abandon it.

  • Run the model comparison before the gap closes, not after. Stand the candidate local model up in your normal environment, point your existing task suite at it, and read the delta while you still have both. Teams that measure afterwards cannot distinguish a model regression from a broken dependency, because inside a gap both present as "it got worse".
  • Re-size the budgets, not just the prompts. If the local model needs more steps for the same task, your step ceilings, timeouts and context budgets were all sized for a different model. The symptom is a mysterious rise in truncated runs, which reads like an infrastructure problem and is not.
  • Switch the cost model from tokens to capacity. There is no per-token bill and no elastic burst inside the gap. You size for peak concurrency and eat the idle GPUs, or you implement admission control and eat the queue. Cost per completed task is still the right metric; the denominator just stops being a vendor invoice.

The dependencies nobody wrote down

Which assets need re-validation in each deployment mode A matrix with five asset rows — prompts and skills, tool schemas, step and timeout budgets, the eval suite, and the cost model — against three deployment modes. Sovereign cloud leaves nearly everything unchanged, self-hosted requires re-checks, and air-gapped requires most assets to be re-earned against a different model. Re-validation surface by deployment mode SOVEREIGN CLOUD SELF-HOSTED AIR-GAPPED Prompts & skills unchanged re-check re-earn Tool schemas unchanged re-check re-check Step & timeout budgets unchanged re-check re-size Eval suite + judge unchanged re-host judge rebuild inside Cost model per-token capacity, not tokens capacity, not tokens unchanged needs a re-check needs rebuilding or re-earning Rows two and three are the quiet ones: a different model changes schema tolerance and how many steps a task takes.
The re-validation surface. Rows two and three are the quiet ones.

A coding agent is a program that looks things up constantly, and most of those lookups were never architecture decisions. Cut the network and they do not raise exceptions — the model answers from its weights instead, which is the worst available failure mode because it is indistinguishable from working.

The version lookup is the canonical case. An agent that could read a package index now states a version number from memory, bounded by a knowledge cutoff it has no way to flag. The tool registry is the same shape: discovery stops, and the agent works from whatever local catalogue was current when it was last refreshed. The hosted judge in your eval pipeline simply errors or silently skips, which is how the environment with the least observability ends up as the one running with no quality gate. Telemetry and error reporting drop into a void. Certificate revocation checks and NTP fail intermittently, producing the 3am incident that looks like nothing.

Every one of those has a fix — a mirrored docs corpus, an internal package index, a refreshed tool catalogue, a locally hosted judge calibrated against the labelled set you already have — and none of them are on a migration checklist written from the product page. The test that finds them costs a day: block egress in a staging environment that is otherwise identical, run the full task suite, and count two things. How many tool calls failed, and how many runs produced a confidently wrong answer instead of failing. The second count is the real cost of the gap.

The one thing the gap gives you back

It is worth being fair about the upside, because it is real and it is underrated: inside a gap, nothing changes under you. No silent endpoint upgrade, no default effort setting moving on a Tuesday, no deprecation window you did not plan for. You own the version — which is the same posture we have argued for elsewhere as pinning and verification, enforced by topology instead of by discipline.

The bill for that arrives as a release process. Model versions, container images, tool servers, dependencies and vulnerability feeds all now cross on a deliberate, authenticated import, and the default outcome of a deliberate process nobody scheduled is that nothing gets updated for eleven months. Put a cadence on the import with a named owner, verify everything inbound by digest and signature — the transfer medium is your supply chain now — and mirror the advisory feeds, because isolation does not patch anything, it only removes your notification channel.

And write the exception path down before you need it. When a production incident needs vendor help and the vendor can see nothing, the improvised answer is someone photographing logs on a phone. Decide in advance what may leave, redacted how, approved by whom.

When to pick which

If your actual requirement is…Sovereign cloudSelf-hosted, hybrid modelAir-gapped, local model
Data stays in a named jurisdictionSufficient, and cheapestOverkillOverkill
Source code never reaches a third-party AI serviceNoNo — prompts still leaveYes, this is the case it answers
Best available model qualityYesYesNo, and the matrix says so
Demonstrable isolation for an auditorContractual onlyPartialYes, and it is the only one that is
Nothing changes without your approvalNoPartlyYes — at the price of a release process

The most common mistake in this table is buying the third row's cost to satisfy the first row's requirement. If a residency clause is what you actually have, stopping at sovereign cloud is a legitimate answer and it preserves the model you built against.

FAQ

Does self-hosted Bob run frontier models at all?

Through hybrid and private-SaaS configurations, yes — that is where Claude Sonnet 5.0 and Opus 4.8, Gemini 3.7 Flash and GPT 5.6 Sol are listed. But a hybrid configuration sends prompts to an external service, which answers a residency requirement and does not answer an isolation one. The models supported on customer-managed infrastructure at general availability are Nemotron and Laguna.

Is this specific to IBM?

No. Any vendor shipping an air-gapped agent product faces the same constraint — you can only run weights the customer may hold — so the pattern of "harness ports, model tier changes" is structural rather than a choice IBM made. IBM is a useful case because its matrix states the split plainly instead of leaving you to infer it.

How much quality do I actually lose?

Unknowable in the abstract and cheap to measure in your own case: run your task suite against the candidate local model before you commit, and compare cost per completed task rather than a benchmark score. The answer is workload-specific, and for well-scoped tasks with good tests it is often much smaller than the leaderboard gap suggests.

Can I keep my existing eval suite?

The task definitions, yes. The judge, usually not — a hosted judge model is an egress dependency, so you re-host a smaller judge inside and re-calibrate it against your existing labelled set, because a different judge is a different metric. Where ground truth can be compared by rule, prefer the rule and remove the dependency entirely.

What is the single best predictor that an air-gap project will go badly?

No one has run the full task suite with egress blocked before the decision. That one experiment surfaces every implicit dependency in an afternoon, and skipping it is why these migrations get discovered rather than planned.

Further reading

On this wiki:

Sources: