Prompt portability.
Changing a model string is a one-line diff that no reviewer can approve, because the thing it changes is not in the diff. A prompt is not a specification of what you want; it is a record of what one model got wrong and how you talked it out of that — so the instruction travels to a new model and the calibration wrapped around it does not. Treat every model swap as a release gated on your own task suite, and write prompts so the model-specific parts are in one place you can throw away.
Four layers in every prompt, and only one of them travels.
It helps to stop treating "the prompt" as one artifact. What you ship is four layers stacked in a single string, and they have wildly different portability.
- Intent — the task, the role, the success criterion. This is the layer people think they are writing, and it is genuinely portable. It is also usually the shortest part of the file.
- Contracts — output shape, schemas, tool signatures. Portable in principle, not in practice: vendors accept different subsets of JSON Schema, so a tool definition that validates on one API is rejected or silently coerced on another. See JSON Schema subsets per vendor.
- Calibration — how hard it tries, how much it hedges, how long it answers, when it asks a clarifying question instead of guessing, where it refuses. This layer is the bulk of a mature prompt and none of it transfers, because every line of it was fitted to one model's behaviour.
- Harness coupling — the chat template, how tool calls are encoded, whether reasoning blocks are returned and must be echoed back, what the effort or thinking default is. This layer is mostly invisible in your source, which is exactly why it surprises you.
The invisible layer moves without you touching anything. Claude Opus 5.5 was the first Claude model to default to medium rather than high effort — so a deployment that swapped the model string and nothing else got a different amount of thinking per step, with the prompt byte-identical. A default is part of your prompt whether or not you wrote it down.
Mature prompts are scar tissue, and scars do not transplant.
Walk back through how a long system prompt got long. Almost every clause was added after a specific failure: "always include the file path" exists because the model dropped paths; "do not apologise" because it opened three replies in a row with an apology; "if the table is empty, say so rather than inventing a row" because once it invented a row. The prompt is a compressed log of one model's failure distribution.
Port that to a different model and both halves of the mismatch cost you. Instructions aimed at failures the new model never had are dead weight — tokens on every request, and occasionally harmful, because an instruction suppressing a behaviour the model does not have will distort a behaviour it does. Meanwhile the new model's actual failures have no clause at all, and you will not discover them by reading the prompt.
- Annotate clauses with the failure that motivated them — a trailing comment, a sibling file, a line in the commit. The annotation is the only thing that makes a port reviewable: it turns "delete this line?" into "is this failure still present?"
- Periodically test clause removal against the model you are on. Prompts accumulate; nothing in a normal workflow ever removes a line, so a two-year-old prompt is carrying instructions for three retired models.
- Expect prompt length to be negatively correlated with portability. The more tuned a prompt is, the more of it is model-specific by construction — this is the uncomfortable corollary of prompt optimization, and it applies with full force to automatically optimized prompts, which are fitted harder than hand-written ones.
Make the model-specific parts separable on purpose.
"Model-agnostic prompt" is not a thing you can write; portability is a property of the surrounding structure, not of careful wording. Four structural moves do nearly all the work.
- Put the quirks in one block. One clearly labelled model-specific section at the end of the system prompt, holding every clause that exists because of this model. Swapping models then means replacing a section, not auditing a wall of text. Everything above it is intent and contracts.
- Replace phrasing with enforcement. "Always output valid JSON" is a prompt-layer wish; a schema plus a validator plus one reprompt is a check that holds on any model, and constrained decoding makes it structural where the provider offers it. Every constraint you move from prose into machinery is a constraint you stop re-tuning.
- Use the API field, not the paragraph. If there is a response-format parameter, a tool schema, or an effort setting, set it explicitly instead of describing the behaviour in prose. Explicit settings survive a model change visibly — they either exist on the new provider or they fail loudly. Prose silently means something slightly different.
- Keep few-shot examples as data. Examples inline in a prose prompt cannot be re-selected per model; examples in a file can, and the right examples for a weaker model are more numerous and more literal than for a stronger one. See few-shot prompting.
Then pick a support matrix and keep it honest: two models you actually test, not "any model". A prompt that has been run against two genuinely different models is far more portable than one written to be universal and run against one, because the second model is what finds the clauses that were never intent in the first place.
The gate on a model swap is an eval run, not a code review.
This is the practical consequence, and it is the one most teams skip. A model change produces a diff of one line, passes review in thirty seconds, and alters behaviour on every request. The only control that matches that shape is a task suite re-run with the old and new model side by side — agent evaluation used as a release gate rather than a research activity.
Compare more than the headline pass rate, because the regressions that hurt are not in the mean.
- Format validity and tool-call validity — the first thing to move, and the easiest to miss if your harness retries on parse failure. A silent 3% rise in retries is a cost and latency regression wearing a success rate.
- Refusal and abstention rate — where a model draws its boundary is part of the product, and it moves between models and between tiers of the same model. See refusals and capability gating.
- Steps and cost per completed task — a model that scores the same and needs 40% more steps is a worse deployment at the same accuracy. Per-token price answers none of this; see agent cost control.
- The tail, not the mean. Look at p95 step count and the distribution of failure modes, not just how many runs passed. A changed failure mode with an unchanged pass rate still invalidates your runbooks.
And treat a deployment change that forces a model change as the same event. When an agent platform moves on-premises or into an air-gapped network, the model tier available there is typically not the one you built against — IBM's self-hosted Bob, generally available on 1 October 2026, keeps the harness (its shell, parallel tool calling, skills, modes) while the models it can run inside the gap are NVIDIA Nemotron and Poolside Laguna rather than the hosted Claude, Gemini and GPT options. The harness ported; the prompts have to be re-earned.
Do this before your next model migration: add a comment to every clause in your system prompt naming the failure it was added for, then delete the ones whose failure you cannot name or reproduce. On most mature prompts this removes a quarter of the text without a measurable quality change — and what is left is a prompt you can actually port, because you now know which lines are about the task and which were about a model you no longer run.
Related: system vs user prompts for where each layer belongs, the agent harness for what sits between your prompt and the model, and model deprecation and migration for running the swap in production.