Interpretability & probes.
Interpretability stopped being a research aesthetic in January 2026, when Google DeepMind put activation probes in front of live traffic on Gemini 2.5 Flash — and the thing that shipped was the least glamorous artifact the field had produced: a linear classifier reading one layer's activations, running at roughly a fiftieth the cost of an LLM classifier and beating it on accuracy. That is the practical shape of interpretability for the next few years. It reads a signal no text filter can see, it is cheap enough to run on every request, it falls apart the moment your traffic drifts, and it needs the weights — which means for most teams it is a question to ask your vendor, not a project to staff.
A probe is a logistic regression with a good view.
Strip the vocabulary away and the mechanism is small. A transformer passes a vector — the residual stream — through each layer. A probe taps that vector at some layer, and a classifier fitted on it predicts a property of the input or the generation in progress. In the overwhelming majority of deployed cases that classifier is linear: a weight vector and a threshold, fitted by logistic regression on a few thousand labelled examples.
- The training set is the whole project. You need examples labelled with the property you want — deceptive vs honest completions, cyber-offensive vs benign requests, on-task vs hijacked. Fitting takes minutes; collecting and labelling the examples takes weeks, and its quality sets your ceiling.
- Which layer matters, and picking one is a mistake. Signal strength varies sharply by depth, and the best layer differs by property and by model. Ensembling several layers recovers performance where any single layer fails, so read a handful and combine them rather than hunting for the one true layer.
- The output is a score per token, not per request. You get a real-valued activation for each token the model emits, which means you can watch a signal rise mid-generation instead of waiting for a finished answer to classify. That is a genuinely different affordance from an output filter.
- Cost is close to zero. A dot product against a vector you already computed, per layer you read. This is why the economics work at all — see the cascade in model routing, which is the same shape applied to quality instead of safety.
"Probe" and "interpretability" get used interchangeably and should not be. Probing asks is this property linearly readable from the activations, which is an engineering question with a benchmark. Mechanistic interpretability asks what computation produces it, which is a science question. The first one ships; the second one is what might eventually make the first one trustworthy.
Why it sees things a text classifier cannot.
Every guardrail you have today reads surface form: the prompt's words, the answer's words, a retrieved document's words. A probe reads the representation the model built before it chose any words. Those come apart in exactly the cases you care about most.
- The model can know something its output does not say. Probes trained to detect strategic deception on simple role-play scenarios generalise to harder multi-turn settings with AU-ROCs reported between 0.96 and 0.999 on held-out evaluation sets. A text classifier reading the same fluent, confident, false paragraph has nothing to go on.
- Paraphrase does not move it much. Surface filters are defeated by rewording, encoding, and translation — the standard injection toolkit. A representation-level signal is a harder target because the model has to understand the request for the request to work, and understanding it is what the probe is reading.
- It fires early. Because the score exists per token, a probe can trip partway through a generation and let you abort before the tokens reach a renderer or a tool. Compare streaming and partial output, where the same mid-generation timing is the source of the problem rather than the fix.
- It is not a lie detector, and this is the important caveat. A probe detects a correlate of the property in the training distribution. Some of the best-known deception-probe results have been shown to lean on textual evidence present in the activations rather than on any privileged internal "intent" — so a probe that scores well can still be reading a proxy you did not intend.
The fragility gap is the number that decides deployment.
Probes hit near-ceiling numbers on benchmarks and then fall over in production, reliably enough that the pattern has a name. The distance between in-distribution and shifted performance is the probe fragility gap, and it is not a rounding error.
- Context length is a shift. DeepMind's own reporting singles this out: probes trained on short-context examples degrade when applied to long-context traffic. An agent's context is long and grows within a run, so this is not an edge case for agentic deployments — it is the normal operating regime.
- Your traffic is a different distribution from the training set, always. The same problem as calibration: a threshold fitted on someone else's data is not a threshold on yours, and nothing tells you when it slipped.
- A model update invalidates the probe outright. The weights the probe reads are gone. This is not a recalibration, it is a refit, and it belongs on the checklist in model deprecation and migration alongside everything else that silently breaks.
- Therefore: cascade, do not gate. The deployed DeepMind design routes roughly 8% of traffic — the probe's uncertain band — to a full LLM classifier, and lands better accuracy than the LLM alone at about 1/50th of its inference cost. The probe is a cheap triage layer that decides what deserves expensive scrutiny. It is a bad sole authority and an excellent first filter.
The cost argument is the whole argument. An LLM classifier on 100% of traffic is usually unaffordable, so teams sample — and a sampled detector misses the attack by design. A probe on 100% plus an LLM on the 8% the probe flags is cheaper than sampling and strictly better coverage. That trade is why this shipped.
What shipped is not what got famous.
The public face of interpretability for the last several years has been sparse autoencoders — decomposing activations into a large dictionary of human-nameable features. That work is genuinely interesting and it is not what is running in production, and the gap is worth understanding before you plan around it.
- On downstream probing tasks, SAEs repeatedly fail to beat simple baselines. The careful case studies — data scarcity, class imbalance, label noise, covariate shift — find that plain probes on raw activations match or exceed SAE-derived features. Where SAEs do show an edge is steering: intervening on a behaviour rather than detecting it.
- There is no ground truth to check against. Nobody can say what concepts a model "really" uses, so an SAE's interpretation cannot be validated directly — only by whether it helps on a task with a score. That is why the field keeps re-running sanity checks against random baselines.
- Read this as sequencing, not dismissal. The deployable artifact today is the boring one. Treat feature dictionaries as a research programme you follow, and linear probes as a control you can budget for this quarter.
You cannot run a probe on an API you only call.
Every technique on this page needs the activations, and activations are behind the weights. This is the fact that decides what your team can actually do, and it splits the world in three.
- If you serve open weights yourself, probes are available to you today, and they are one of the few genuine capabilities that self-hosting buys beyond price and privacy. Add it to the ledger in open-weight vs closed models, which usually gets argued on cost and licensing alone.
- If you call a hosted frontier model, you get whatever probe-derived protections the provider chose to run, described — at best — in the system card. You cannot fit one for your own threat model, your own tenants, or your own definition of off-task. Asking a vendor whether internal-state monitoring runs on your traffic, and whether its outcome is visible to you, is a sharper procurement question than most security questionnaires contain.
- If you need this and cannot self-host the main model, the workable pattern is a small open model running alongside as the monitored surface — the redaction-and-classification role in small and local models. It probes its own activations, not the frontier model's, so it catches what it can see and nothing more. Be honest about that limit.
Do not start here. A probe is a strong third control and a terrible first one: if untrusted text can still reach a dangerous tool, an 0.99 AU-ROC detector is the wrong purchase, because the fix is to cut the path rather than to notice traffic on it. Build the deterministic controls first — scoped credentials, egress limits, approval gates on consequential actions. Then, if you run the weights and you have a labelled corpus of your own failures, fit a probe on it, ensemble a few layers, deploy it as the cheap tier of a cascade with an LLM judge behind it, and re-fit it on every model change. If you do not run the weights, spend the same effort on behavioural detection, which reads the trajectory instead of the activations and is available to everyone.
Related: chain-of-thought faithfulness for why the reasoning text is not a window into the computation, guardrails for the layer a probe joins, and evaluating guardrails and detectors for how to measure one honestly before you trust it.