The jagged frontier.
A model that drafts a competent legal summary can fail at counting the items in the list it just wrote, and from the outside those two tasks look equally easy. That is the jagged frontier: capability is not a level, it is a coastline with inlets, and nothing about performance on one task predicts performance on the task next to it. The practical consequence is that a benchmark score, a demo, and a vendor claim are all evidence about somewhere else — the only measurement that transfers to your task is a measurement on your task.
Where the term comes from, and the number that made it stick.
The phrase comes from a pre-registered field experiment run with 758 Boston Consulting Group consultants, published as Navigating the Jagged Technological Frontier. On 18 realistic consulting tasks chosen to sit inside the model's capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at more than 40% higher quality. On one task deliberately built to sit just outside it, consultants using AI were 19 percentage points less likely to reach a correct answer than consultants with no AI at all.
Two things make that pair of results worth carrying around rather than filing under "AI has limits".
- The inside-tasks result is large. This is not a marginal tool. Where the frontier includes your task, the effect size is the kind organisations restructure around.
- The outside-task result is negative, not merely absent. Assistance did not fail to help; it actively made people worse than working alone, because a fluent wrong answer is more persuasive than no answer. The cost of being outside the frontier is paid by the human who trusted the output.
And the tasks were not separable in advance by any ordinary notion of difficulty. That is the whole point: the boundary does not follow the contour lines a person would draw.
The edge is invisible because failure keeps the shape of success.
Human competence degrades legibly. A person working past their ability slows down, hedges, asks a question, or produces something visibly rough — and the effort is itself a signal to the person watching. A model has no such tell. Just outside the frontier it produces output with the same fluency, the same structure, the same confident register as it does well inside, because fluency and correctness are produced by different parts of the process. This is the same asymmetry that hallucination and calibration describe from the inside; the jagged frontier is what it looks like from the outside, at the level of task selection.
So the danger zone is not deep outside the frontier, where the model refuses or the output is obviously wrong. It is the strip just beyond the edge, where the answer is well-formed, plausible, and incorrect — and where the reviewer's attention is lowest, because the previous nine outputs were fine.
This is why "the model is good at X" is a statement with almost no operational content. Ask instead: good at which specific instances of X, checked how, and what did the failures look like? A capability claim without a failure description is a marketing claim, whoever is making it.
What actually puts a task outside, and why agents make it worse.
The frontier is not drawn by difficulty. The properties that put a task outside it are mostly structural:
- The answer depends on something not in the context — a fact after the cutoff, a private document, the current state of a system. The model will supply the missing piece from its own distribution rather than stop.
- The correct output is a refusal or an admission of uncertainty. "There is no answer from this data" is a shape models are weakly trained to produce and strongly pressured against.
- The task is a rare combination of common parts. Each ingredient is familiar; the specific combination is not, and the model interpolates toward the familiar version instead of the one you asked for.
- Success requires holding a constraint across the whole output — a word limit, a schema, an invariant that must be true in paragraph nine because of what paragraph two said.
Then an agent takes that per-step jaggedness and runs it forty times. A loop does not average out the rough edges; it compounds them, because each step's output becomes the next step's input, and one excursion past the frontier at step six is silently laundered into the premise of steps seven through forty. Task horizon is the measurement of how far that goes before it breaks, and trajectories are where you find out which step it was.
Map your own coastline, and re-draw it on every model change.
Because the frontier is jagged, the useful unit of evidence is small and specific. Three habits follow.
- Pilot per task, not per model. The question "should we use AI for this" has a different answer for each of the eleven things your team does, and the answers are not ordered by how hard those things feel. Take twenty real instances of one task, run them, and grade them against what actually counts as correct. This is evaluation at its least glamorous and most decisive.
- Prefer tasks where you can check the answer. If verifying is much cheaper than producing, a task just outside the frontier costs you a retry; if verifying is as expensive as producing, it costs you a wrong decision. That asymmetry is the generator–verifier gap, and it is the single best predictor of whether a deployment survives contact with reality.
- Treat your map as perishable. Every model release redraws the coastline — mostly outward, but not uniformly, and behaviour you depended on can change in both directions. A map drawn against last year's model, never re-run, is the reason teams are still avoiding tasks that became easy and still trusting ones that quietly did not. Eval-set maintenance is how the map stays alive.
Pick the task your team is most confident about and the task it is most sceptical about, and measure both on twenty real cases this week. Expect to be wrong about one of them — that experience, once, does more for an organisation's judgement than any amount of reading about capabilities, because it replaces a belief about "the model" with a number about a task. Related: reading benchmarks critically for why the published scores were never about you, when to use an agent for the prior question, and agent evaluation for doing this at more than twenty cases.