Scaling Laws

F20
Concepts · AI Foundations

Scaling laws.

The most famous scaling law was revised in 2022 by a factor of more than ten, and the revision was not a rounding error — it was the discovery that everyone had been training models far too large on far too little data. That is the thing to carry: scaling laws are an empirical curve fitted to pretraining loss, they are startlingly accurate about that one quantity, and they say nothing whatsoever about whether your agent finishes a forty-step task. Almost every bad inference drawn from them comes from confusing the two.

STEP 1

What the law actually claims.

A scaling law is a fitted power relationship: as you increase compute, parameters, or training data, the model's loss — its average surprise at the next token in held-out text — falls along a smooth, predictable curve. There is no theory behind it. It is a regression through measurements, and its authority comes entirely from the fact that it has held over many orders of magnitude.

  • The quantity being predicted is cross-entropy loss on the pretraining distribution. Not accuracy, not capability, not usefulness. Loss. See what an LLM is for why that number is the training objective in the first place.
  • The practical use is budget allocation. Given a fixed compute budget, the curve tells you how to split it between a bigger model and more tokens. That is a real and valuable answer to a real question — and it is a narrower question than the one people ask the law to settle.
  • Within its domain it is remarkably good. Labs routinely predict a frontier run's final loss from small-scale ablations, before spending the money. That track record is what gives the whole idea its reputation.
STEP 2

The revision that should calibrate your confidence.

In 2020, Kaplan et al. fitted the first widely used version and concluded that as compute grows, parameters should grow faster than data. GPT-3 was built on that advice: 175 billion parameters trained on roughly 300 billion tokens, about 1.7 tokens per parameter.

In 2022, the Chinchilla paper refitted the same relationship with models actually trained to convergence and with learning-rate schedules tuned per model size — two methodological corrections the original had not made. The conclusion inverted. Compute-optimal turned out to be roughly 20 tokens per parameter, and a 70-billion-parameter model trained on 1.4 trillion tokens beat a 280-billion-parameter model trained the old way.

Read that as an epistemics lesson rather than a trivia item. A curve that an entire industry treated as a law was off by more than an order of magnitude in its central prescription, for two years, because of under-tuned learning rates and models that were never trained to the end. Scaling laws are measurements, and measurements have methodology. Treat any extrapolation past the range that was actually measured as a forecast, not a fact.

STEP 3

Loss is not capability, and the gap is where agents live.

The curve is continuous and your task is not. A model either produces a valid tool call or it does not; it either recovers from a failed step or it loops. Those are thresholds, and a smooth decline in loss crosses them at points nobody can predict from the curve.

  • The mapping from loss to task success is not derivable. It is measured, after the fact, per task. This is why a model that is clearly better on paper can be worse in your harness, and why the only honest answer to "will the next model fix this?" is an eval run.
  • Agentic failure is compounding, not average. A per-step reliability of 98% is excellent and gives you a 45% success rate over forty steps. Scaling laws describe average next-token behaviour; task horizon describes what that average does when it is multiplied by itself several dozen times, and it is the more useful curve for anyone building agents.
  • Benchmarks are not the pretraining distribution. Scores on a public benchmark move for reasons that include contamination, formatting, and harness quality — see reading benchmarks. A loss curve is a cleaner measurement precisely because nobody optimises against it directly.
  • "Emergence" is contested and you do not need to resolve it. Whether a sharp jump in a benchmark score reflects a real phase change or a discontinuous metric applied to smooth underlying improvement is a live argument. Either way, the practical consequence is the same: you cannot read your capability off the loss curve.
STEP 4

The curve everyone quotes is no longer the curve that sets your bill.

Pretraining scaling still works. It has simply stopped being where the marginal dollar goes, and the substitutes have their own curves with different shapes — which is why parameter count is now close to useless as a proxy for anything you care about.

  • Post-training moved the frontier for agents. Tool use, instruction following and long-horizon behaviour improved far more from post-training than from another pretraining doubling. None of that is described by the pretraining law.
  • Inference-time compute is a second axis entirely. A reasoning model buys quality with tokens spent per request rather than parameters baked in at training time, and it has its own diminishing-returns curve — see inference-time scaling. This is the axis your bill actually rides on, because you pay it on every single call.
  • Sparsity broke the parameter-to-cost link. A mixture-of-experts model with 400 billion total parameters can cost less per token than a dense 70-billion one, because compute tracks the active parameters and memory tracks the total. "How big is it?" has stopped being one question.
  • Distillation runs the curve backwards. A small model trained on a large one's outputs lands well off the compute-optimal frontier and is the right choice anyway, because your constraint is latency or unit cost rather than training efficiency. See distillation and quantization and small and local models.

Use scaling laws for the one thing they are for — allocating a training budget — and refuse to use them for anything else. When someone argues that a capability is coming because the trend line says so, ask which quantity the trend line plots; if the answer is loss, the argument has not yet touched your problem. In your own work, delete parameter count from your model-selection criteria entirely and replace it with two measurements you take yourself: success rate on your task at your horizon, and cost per successfully completed task. Related: choosing a model, training vs inference, and price deflation and cost per task for what the curve has and has not done to what you pay.