Deep-Dives / Reasoning & Test-Time Compute

Reasoning & Test-Time Compute

Chain of thought, self-consistency, tree/graph of thought, and the inference-time-scaling laws that govern them.

  1. Chain-of-Thought, Properly
    What CoT actually buys (serial compute, not introspection), faithfulness vs post-hoc rationalization, when it hurts, and structured vs free traces.
  2. Self-Consistency & Sampling
    Why sampling + majority vote works, the exact bias-amplification failure, the saturating returns curve, and how to spend the k budget.
  3. Tree & Graph of Thought
    Deliberate search over partial solutions, the multiplicative cost, and the load-bearing dependency on a partial-state scorer.
  4. Verifier-Guided Search
    Outcome vs process reward models steering best-of-N and beam search, reward hacking at inference time, and why the verifier is the product.
  5. Inference-Time Scaling
    Test-time compute as a second scaling axis, the difficulty-adaptive compute-optimal frontier, and where more thinking stops paying.
  6. When Reasoning Helps (and When It Burns Money)
    The synthesis decision rule — task class × verifiability × budget — an escalation ladder, the named money-burning patterns, and a do/don't list.
  7. Adaptive Thinking & Effort Budgets
    `budget_tokens` is deprecated — Claude's `effort`, Gemini's `thinking_level`, OpenAI's `reasoning_effort`, and when the model overrides your budget.
  8. Carrying Reasoning Across Tool Calls
    Reasoning became a signed, opaque input you must hand back unchanged — so injecting a reminder, adding a tool mid-run or trimming history now invalidates it, silently or with a 400. Echo, never rebuild.
  9. Latent Reasoning & the Trace You Stop Getting
    Continuous thought and recurrent depth buy real capability — a two-layer transformer with D continuous steps solves graph reachability where discrete CoT needs O(n²) decoding steps, and a 3.5B recurrent model reaches ~50B-equivalent compute at 50 loops — by deleting the artifact five production controls consume. Name those five, concede that the trace was never faithful and keep it anyway because monitoring it still beats monitoring actions alone, then fit probes on the early latent steps where the signal concentrates — which makes this a procurement question, because a hosted latent model gives you no residual stream to read.