Read IBM's two lists next to each other and the sovereignty trade stops being abstract. Self-hosted Bob went generally available on 1 October 2026 for on-premises, private-cloud, sovereign-cloud and air-gapped environments, and it keeps the agent harness whole — the shell, parallel tool calling, skills, operating modes. The models supported on customer-managed infrastructure at general availability are NVIDIA Nemotron and Poolside Laguna; the hosted and hybrid configurations are where Claude Sonnet 5.0 and Opus 4.8, Gemini 3.7 Flash and GPT 5.6 Sol appear. The portable artifact is the one your migration checklist covers. The artifact that decides whether the output is any good is the one that stayed behind — which means the air-gap project you costed as infrastructure is a re-validation project, and the thing you have to rebuild is your eval suite.
At a glance
Three postures get called the same thing in procurement documents, and they have different prices.
| Posture | Where the code sits | Model you can run | What you must re-validate |
|---|---|---|---|
| Sovereign cloud | Provider infrastructure, contracted region and jurisdiction | Hosted frontier models, unchanged | Essentially nothing technical |
| Self-hosted, hybrid model | Your infrastructure; prompts still leave for inference | Hosted frontier models via an external service | Egress policy, latency budgets, failure handling |
| Self-hosted or air-gapped, local model | Code, development context and build artifacts stay inside | Customer-managed open-weight models — for Bob at GA, Nemotron and Laguna | Prompts, skills, step budgets, the eval suite, the cost model |
The feature list ported; the behaviour did not
IBM's framing is accurate and it is also the source of the confusion. Bob is an agentic software-development platform — understanding code, planning work, executing changes, validating results — and the self-hosted edition genuinely keeps the capabilities that make it that: the shell, the IDE experience, parallel tool calling, skills, modes. If you listed what your developers interact with, almost all of it crosses the gap.
But none of those capabilities decide whether a patch is correct. The harness decides what the agent can attempt; the model decides how often the attempt lands. Swap the model underneath an identical harness and you have changed the only variable that the user experience is actually a function of, while every artifact that mediates between you and the model — system prompts, skills, tool descriptions, step ceilings, retry policies, acceptance thresholds — stays pinned to a model you are no longer running.
Those artifacts are not configuration. They are measurements. A mature agent prompt is a compressed log of one model's failure distribution: this clause exists because that model dropped file paths, this one because it over-edited, this threshold because its confidence was miscalibrated in a particular direction. Carry them across and you pay twice — tokens spent suppressing failures the new model does not have, and no clause at all for the failures it does. That is the general problem in prompt portability; an air gap is just the version of it you cannot postpone.
Read the model matrix as the price list
The practical reading advice for any vendor's sovereign or air-gapped offering is to skip the capability page and go to the supported-model matrix, because that table is the specification. If the models certified for customer-managed infrastructure differ from the hosted list — and at Bob's general availability they do — then the vendor has told you, precisely and in public, which model tier your isolation requirement costs you.
This is not a criticism of the arrangement; it is the only honest way to ship it. An air-gapped deployment can only run weights the customer is licensed to hold and has hardware to serve, which rules out every model whose weights do not leave the lab. NVIDIA Nemotron and Poolside Laguna are in that list because they can be, and a vendor that promised frontier-hosted quality inside a gap would be promising something nobody can deliver.
What follows is a reframing of the project plan rather than a reason to abandon it.
- Run the model comparison before the gap closes, not after. Stand the candidate local model up in your normal environment, point your existing task suite at it, and read the delta while you still have both. Teams that measure afterwards cannot distinguish a model regression from a broken dependency, because inside a gap both present as "it got worse".
- Re-size the budgets, not just the prompts. If the local model needs more steps for the same task, your step ceilings, timeouts and context budgets were all sized for a different model. The symptom is a mysterious rise in truncated runs, which reads like an infrastructure problem and is not.
- Switch the cost model from tokens to capacity. There is no per-token bill and no elastic burst inside the gap. You size for peak concurrency and eat the idle GPUs, or you implement admission control and eat the queue. Cost per completed task is still the right metric; the denominator just stops being a vendor invoice.
The dependencies nobody wrote down
A coding agent is a program that looks things up constantly, and most of those lookups were never architecture decisions. Cut the network and they do not raise exceptions — the model answers from its weights instead, which is the worst available failure mode because it is indistinguishable from working.
The version lookup is the canonical case. An agent that could read a package index now states a version number from memory, bounded by a knowledge cutoff it has no way to flag. The tool registry is the same shape: discovery stops, and the agent works from whatever local catalogue was current when it was last refreshed. The hosted judge in your eval pipeline simply errors or silently skips, which is how the environment with the least observability ends up as the one running with no quality gate. Telemetry and error reporting drop into a void. Certificate revocation checks and NTP fail intermittently, producing the 3am incident that looks like nothing.
Every one of those has a fix — a mirrored docs corpus, an internal package index, a refreshed tool catalogue, a locally hosted judge calibrated against the labelled set you already have — and none of them are on a migration checklist written from the product page. The test that finds them costs a day: block egress in a staging environment that is otherwise identical, run the full task suite, and count two things. How many tool calls failed, and how many runs produced a confidently wrong answer instead of failing. The second count is the real cost of the gap.
The one thing the gap gives you back
It is worth being fair about the upside, because it is real and it is underrated: inside a gap, nothing changes under you. No silent endpoint upgrade, no default effort setting moving on a Tuesday, no deprecation window you did not plan for. You own the version — which is the same posture we have argued for elsewhere as pinning and verification, enforced by topology instead of by discipline.
The bill for that arrives as a release process. Model versions, container images, tool servers, dependencies and vulnerability feeds all now cross on a deliberate, authenticated import, and the default outcome of a deliberate process nobody scheduled is that nothing gets updated for eleven months. Put a cadence on the import with a named owner, verify everything inbound by digest and signature — the transfer medium is your supply chain now — and mirror the advisory feeds, because isolation does not patch anything, it only removes your notification channel.
And write the exception path down before you need it. When a production incident needs vendor help and the vendor can see nothing, the improvised answer is someone photographing logs on a phone. Decide in advance what may leave, redacted how, approved by whom.
When to pick which
| If your actual requirement is… | Sovereign cloud | Self-hosted, hybrid model | Air-gapped, local model |
|---|---|---|---|
| Data stays in a named jurisdiction | Sufficient, and cheapest | Overkill | Overkill |
| Source code never reaches a third-party AI service | No | No — prompts still leave | Yes, this is the case it answers |
| Best available model quality | Yes | Yes | No, and the matrix says so |
| Demonstrable isolation for an auditor | Contractual only | Partial | Yes, and it is the only one that is |
| Nothing changes without your approval | No | Partly | Yes — at the price of a release process |
The most common mistake in this table is buying the third row's cost to satisfy the first row's requirement. If a residency clause is what you actually have, stopping at sovereign cloud is a legitimate answer and it preserves the model you built against.
FAQ
Does self-hosted Bob run frontier models at all?
Through hybrid and private-SaaS configurations, yes — that is where Claude Sonnet 5.0 and Opus 4.8, Gemini 3.7 Flash and GPT 5.6 Sol are listed. But a hybrid configuration sends prompts to an external service, which answers a residency requirement and does not answer an isolation one. The models supported on customer-managed infrastructure at general availability are Nemotron and Laguna.
Is this specific to IBM?
No. Any vendor shipping an air-gapped agent product faces the same constraint — you can only run weights the customer may hold — so the pattern of "harness ports, model tier changes" is structural rather than a choice IBM made. IBM is a useful case because its matrix states the split plainly instead of leaving you to infer it.
How much quality do I actually lose?
Unknowable in the abstract and cheap to measure in your own case: run your task suite against the candidate local model before you commit, and compare cost per completed task rather than a benchmark score. The answer is workload-specific, and for well-scoped tasks with good tests it is often much smaller than the leaderboard gap suggests.
Can I keep my existing eval suite?
The task definitions, yes. The judge, usually not — a hosted judge model is an egress dependency, so you re-host a smaller judge inside and re-calibrate it against your existing labelled set, because a different judge is a different metric. Where ground truth can be compared by rule, prefer the rule and remove the dependency entirely.
What is the single best predictor that an air-gap project will go badly?
No one has run the full task suite with egress blocked before the decision. That one experiment surfaces every implicit dependency in an afternoon, and skipping it is why these migrations get discovered rather than planned.
Further reading
On this wiki:
- Air-gapped agent deployments — the operational playbook, including the import cadence and the acceptance gate.
- Prompt portability — why the prompts are the asset that does not travel.
- Self-hosted inference for agents — serving the model you can actually run.
- Data residency and sovereignty — which of the three postures your requirement really needs.
- Egress control for agents — the policy half, for deployments that are not fully isolated.
Sources:
- IBM Bob expands to self-hosted environments — IBM announcement
- IBM allows on-prem deployment of its Bob agentic development platform — SiliconANGLE, 1 October 2026