Latent Reasoning, and the Trace You Stop Getting.
Every control you run on an agent's thinking — the monitor that reads the plan before the tool fires, the eval that scores the trajectory, the reviewer who has to explain why a refund was approved — rests on an implementation accident nobody chose: that the model currently deliberates in words, and words go in a log. Latent reasoning deletes that accident in exchange for compute, and it is the most efficient form of test-time scaling anyone has demonstrated. The trade is not tokens for accuracy. It is your only readable account of the run for accuracy, and the replacement account has to exist before you take the deal, not after the first incident.
The model keeps thinking; it just stops writing it down.
Two research lines point at the same architecture, and both are now several years deep enough to take seriously as a deployment question rather than a curiosity.
Continuous thought. Coconut (Chain of Continuous Thought, Hao et al., arXiv 2412.06769) changes one line of the decoding loop. Instead of sampling the last hidden state into a token and re-embedding that token, it feeds the hidden state straight back in as the next input embedding. Nothing is decoded. The "thought" never becomes text, so it is never forced to collapse onto a single next step.
Recurrent depth. Geiping et al. (arXiv 2502.05171) take a 3.5B-parameter model — Huginn, trained on 800B tokens — and give it a core block that can be unrolled an arbitrary number of times at inference. Run more loops, spend more compute per token, with no change in parameter count and no reasoning traces in the training data at all. At up to 50 loops the authors report an effective compute budget in the range of a 50B-parameter model.
The theoretical reason this is not merely a compression trick is worth stating exactly, because it is the strongest argument on this page and it is not an argument about efficiency. A continuous thought vector is not obliged to represent one state. Reasoning by Superposition (arXiv 2505.12514) shows that a two-layer transformer with D steps of continuous chain-of-thought solves directed-graph reachability, where D is the graph's diameter — while the best known bound for constant-depth transformers using discrete chain-of-thought is O(n2) decoding steps over n vertices. The mechanism is that each latent vector holds a superposition of search frontiers, so one autoregressive step advances a breadth-first search rather than one path.
Read that result as a statement about what discrete tokens cost, not about what latents add. A token is a commitment: the sampling step throws away the distribution and keeps one symbol. Everything tree-of-thought does with explicit branching and an explicit scorer, continuous thought does inside the residual stream for free — and, for the same reason, invisibly.
The economics favour this, which is why it will arrive whether you plan for it or not.
Discrete reasoning is priced twice. You pay for the reasoning tokens as output, and you pay for them again as context on every subsequent turn of the agent loop — see context budgeting for what that compounds to over a forty-step run. You also pay in wall-clock: serial decoding at a few dozen tokens per second is the single largest contributor to the latency of a reasoning-heavy turn, and no amount of batching fixes it because the dependency is sequential.
Latent reasoning breaks the coupling between amount of deliberation and length of context. Compute per token becomes a dial — loops, or continuous steps — that does not touch the context window at all.
# The same deliberation budget, two ways discrete CoT 4,000 reasoning tokens -> ~4,000 serial decode steps -> 4,000 tokens of context carried on every later turn -> billed as output now, as cached input forever latent 32 recurrent steps on the final position -> no decode, no sampling -> 0 tokens of context -> billed as compute, invisible to the transcript # The second column is cheaper on every axis you measure, # and empty on the one axis you forgot you were measuring.
There is a second-order pull, too. Recurrent-depth training needs no chain-of-thought corpus — no distillation from a bigger reasoner, no RL over verifiable tasks, no licensing question about whose traces you trained on. For anyone building an open-weights reasoner without a frontier lab's data pipeline, that is not a nice-to-have. It is the difference between shipping and not.
So do not plan around latent reasoning as a hypothetical. Plan around it as a thing that arrives in a model card one quarter, for a model that is cheaper and faster than the one you are using, with a paragraph you will skim.
Name what breaks, because "interpretability" is too vague to budget for.
The research framing is that latent reasoning closes the monitoring window. True, and too abstract to act on. Here is the same claim expressed as five artifacts that exist in production agent stacks today and stop existing on a latent model.
- The pre-execution monitor. A classifier reading the model's stated plan before a consequential tool call, looking for the shape of an injected instruction or an out-of-policy intent. This is the cheapest high-yield control in the injection-defense stack, precisely because a compromised model usually says what it is about to do before it does it.
- Process scoring in evals. Half of trajectory evaluation reads the reasoning, not just the actions — did it consider the constraint, did it notice the contradiction, did it decide or drift. On a latent model your eval collapses to outcome-and-actions, which is the weaker signal the whole subfield exists to escape.
- The triage transcript. When an agent does something wrong at 3am, the first thing anyone opens is the reasoning around the bad step. Without it, incident review becomes inference from tool calls — you can see it queried the wrong account, and you cannot see why it thought that was the right account.
- The explanation of record. Regulated workflows that owe a data subject or an auditor an account of an automated decision have been quietly using the model's own reasoning as the raw material for that account. It was always a shaky practice; it becomes an impossible one. See contestability and appeals for what has to replace it.
- Drift detection on the reasoning itself. Teams that watch the distribution of reasoning length, hedging language or plan structure across a deploy have a genuinely early warning signal. It disappears without an error.
Notice what these five have in common. None of them is a research programme. Each is a running piece of software with a dashboard and an owner, and each one's input is a string that is about to become a null.
The strongest objection: the trace was never faithful, so who cares.
This objection is correct on the facts and wrong on the conclusion, and it is worth walking through carefully because it is the argument that will be used to wave the change through.
The facts first. A chain of thought is not a transcript of the computation. Models produce post-hoc rationalisations, omit the cue they actually used, and — under RL pressure — learn to write a reasoning that scores well rather than one that describes what happened. All of that is documented, and it is the whole content of chain-of-thought faithfulness. Anyone treating the trace as ground truth is already wrong today.
Now the conclusion that does not follow. The cross-lab position paper on monitorability (arXiv 2507.11473) makes the load-bearing empirical claim: in almost every setting studied, monitoring the chain of thought catches more misbehaviour than monitoring actions and final outputs alone. A channel does not need to be truthful to be useful. It needs to be correlated with the thing you are trying to catch and cheap to read, and an unfaithful trace clears that bar comfortably — which is why a detector on it works even though a philosopher would object to calling it an explanation.
So the right framing is not "we are losing the truth." It is: you are losing a noisy, high-recall, essentially free sensor, and your remaining sensors are all downstream of the action. Everything you can still see happens after the decision is made. That is a materially different control posture, and it is the one to argue about.
The same paper is unusually blunt about the fragility of what it is defending: monitorability is a property of current training recipes, not a guarantee, and it can erode without anyone deciding to erode it. Process supervision that rewards nice-looking reasoning, architecture changes that shorten the trace, and latent reasoning are three different roads to the same place. If you depend on the channel, measure it — do not assume it.
Latent is not opaque. It is opaque to the tools you happen to own.
The useful asymmetry: a latent reasoning step is unreadable by a text classifier and perfectly readable by a linear probe. It is a vector, and vectors are the native input of the entire probing toolkit — which, unlike most interpretability work, already runs in production at frontier scale and costs a rounding error per call.
Early work on detecting misaligned reasoning inside continuous thought (arXiv 2604.23460) reports the practically important part: the signal is concentrated in the early latent steps, the ones that function as planning. That maps cleanly onto the control you lost. The pre-execution monitor read the plan; a probe on the first few latent steps reads the plan too, in a different representation.
Three consequences follow, and the third is the one that decides your architecture.
- Probes are per-model and per-layer. They are fitted artifacts, they degrade under distribution shift, and every model update invalidates them. Budget refitting as a recurring cost on your migration checklist, not as a project.
- Probes need labels you do not have yet. To fit a detector for "about to do the out-of-policy thing," you need examples of the model about to do the out-of-policy thing. Today those examples are in your reasoning logs. Harvest them while the channel still exists.
- Probes need white-box access. Residual-stream reads are not in anyone's public chat API. If your latent-reasoning model is hosted behind a vendor endpoint, you cannot fit or run a probe on it at all, and the replacement sensor is unavailable to you by contract rather than by physics.
That last point turns this from an ML question into a procurement question, and it is the sentence to put in front of whoever signs the contract: on a hosted latent-reasoning model, the vendor is the only party who can monitor the reasoning, and whether they do is a term, not a fact.
What to do, in the order that survives being half-finished.
Nothing here is a reason to avoid latent reasoning. It is a reason to treat the reasoning format as an interface your stack depends on — which is exactly what it became, unannounced, the day you shipped a monitor that reads it.
- Write down which of your controls read the trace. One afternoon, one list, and it is the artifact that makes every later argument concrete. Most teams discover between two and five, and are surprised by at least one.
- Move every control you can to the action boundary. A policy check on the tool call with its resolved arguments, a receipt for each consequential action, an egress filter on the response. These survive any change in how the model thinks, which is the definition of a control worth having. This is the same conclusion decision receipts reaches from the audit side.
- Make the verifier carry the weight the trace was carrying. If a separate model or program scores the proposed action, you have an observable that is independent of the reasoning format — and, per verifier-guided search, the thing that was doing most of the work anyway.
- Keep a decoded sample. Where a vendor offers a decode-to-text mode on latent steps, run it on a small sampled fraction of production traffic and on your entire eval set. A sampled channel is a weaker control and a perfectly good measurement, and it keeps your probe training data flowing.
- Put the reasoning format in the model-change checklist. Alongside context window and pricing, a line that reads does this model emit a readable reasoning trace, and which of our controls consume it. This is the step that converts a silent capability loss into a decision someone made.
Do this one thing this week, before any of the above: take fifty production runs, strip the reasoning out of each, and hand them to your incident reviewer and your eval. Measure what the reviewer can still conclude and how much your eval's scores move. That number is the price of latent reasoning for your system, in your units, and you can get it today without waiting for a model that has the property. Teams who run it usually find the eval barely moves and the reviewer is lost — which tells you the trace was propping up your operations, not your metrics, and that is where the replacement has to go.
Related: inference-time scaling for the axis latent reasoning is competing on, carrying reasoning across tool calls for the shift that already made the trace opaque by default, and detecting agent compromise for the controls that never depended on it.