Deep-Dives / Reasoning & Test-Time Compute
Reasoning & Test-Time Compute
Chain of thought, self-consistency, tree/graph of thought, and the inference-time-scaling laws that govern them.
- Chain-of-Thought, ProperlyWhat CoT actually buys (serial compute, not introspection), faithfulness vs post-hoc rationalization, when it hurts, and structured vs free traces.
- Self-Consistency & SamplingWhy sampling + majority vote works, the exact bias-amplification failure, the saturating returns curve, and how to spend the k budget.
- Tree & Graph of ThoughtDeliberate search over partial solutions, the multiplicative cost, and the load-bearing dependency on a partial-state scorer.
- Verifier-Guided SearchOutcome vs process reward models steering best-of-N and beam search, reward hacking at inference time, and why the verifier is the product.
- Inference-Time ScalingTest-time compute as a second scaling axis, the difficulty-adaptive compute-optimal frontier, and where more thinking stops paying.
- When Reasoning Helps (and When It Burns Money)The synthesis decision rule — task class × verifiability × budget — an escalation ladder, the named money-burning patterns, and a do/don't list.
- Adaptive Thinking & Effort Budgets`budget_tokens` is deprecated — Claude's `effort`, Gemini's `thinking_level`, OpenAI's `reasoning_effort`, and when the model overrides your budget.
- Carrying Reasoning Across Tool CallsReasoning became a signed, opaque input you must hand back unchanged — so injecting a reminder, adding a tool mid-run or trimming history now invalidates it, silently or with a 400. Echo, never rebuild.
- Latent Reasoning & the Trace You Stop GettingContinuous thought and recurrent depth buy real capability — a two-layer transformer with D continuous steps solves graph reachability where discrete CoT needs O(n²) decoding steps, and a 3.5B recurrent model reaches ~50B-equivalent compute at 50 loops — by deleting the artifact five production controls consume. Name those five, concede that the trace was never faithful and keep it anyway because monitoring it still beats monitoring actions alone, then fit probes on the early latent steps where the signal concentrates — which makes this a procurement question, because a hosted latent model gives you no residual stream to read.