A score without a human baseline has no denominator.
"62% on our task set" is not a measurement of anything until you know what a competent person scores on the same tasks, under the same time limit, with the same tools, graded by the same grader — and for almost every benchmark an agent team quotes, that number was never collected. The consequence is not academic: teams ship agents whose measured performance is above a baseline nobody ran and below one nobody ran either, then argue about a gap whose sign is unknown. A human baseline is not a constant you can look up. It is a four-part tuple, and changing any part of it moves the number further than a model generation does.
What the phrase has to mean before it means anything.
"Human-level" is used as though humans had a level. They do not — they have a performance surface over conditions, and a baseline is a single point on it. Four coordinates fix the point, and a baseline that does not state all four is not comparable to anything, including a later run of itself.
# A human baseline is this tuple, or it is a rumour who expertise, domain familiarity, familiarity with THIS codebase budget wall-clock allowed per task, and whether it was enforced tools search? a debugger? the internet? an LLM? grader the identical scorer, applied to human and agent output # Report format that survives a year human 41/60 @ 2h cap, n=6 annotators, same rubric agent 37/60 @ 2h cap, same rubric, pass^3 ratio 0.90 (CI overlaps; see paired test)
The fourth coordinate is the one almost everyone drops, and it is the one that inverts conclusions. If the agent is graded by an automated checker and the human by a reviewer reading the diff, you have not compared two workers on one task; you have compared two grading regimes on two workers. The trajectory and process evaluation question — what exactly are we scoring — has to be answered once and then applied to both sides without adjustment.
The practical test for whether you have a baseline: could a stranger reproduce your human number to within a few points from your write-up alone? If the write-up says "human experts achieve roughly 90%" and cites a paper that measured different people on a different time budget, you have imported a constant, not a baseline.
The coordinates dominate the model, and the time budget dominates the coordinates.
It is easy to assume the tuple is a second-order correction on a first-order fact about capability. It is the other way round on realistic agent tasks, because the quantity being measured — did a multi-step piece of work reach a verifiable end state — is extremely sensitive to how long the worker had and what they already knew.
- Time budget. Human success on a multi-hour software task is roughly monotone in time allowed, over a wide range, and the range is wide precisely where interesting tasks live. An uncapped human baseline and a step-limited agent are not on the same axis, and the comparison flatters whichever side you forgot to cap. Agents are almost always capped, by context, step limits and cost — see retry amplification for how those caps interact.
- Familiarity. On repository-grounded tasks, a contributor who knows the codebase and an equally skilled engineer who has never seen it are not the same baseline, and the gap is large. Which one you hire for the baseline decides whether your agent looks close to expert or close to superhuman.
- Tool access. A human baseline collected without internet access in 2023 is measuring a worker your agent is not competing with. A 2026 baseline where annotators may use an assistant is measuring a human-plus-agent system, which is usually the right comparison for a deployment decision and the wrong one for a capability claim. Pick deliberately and label it.
- Variance between annotators. Individual completion times on the same task differ by multiples, not percentages. METR aggregates multiple successful baselines with the geometric mean for exactly this reason — the arithmetic mean of a heavy-tailed duration distribution is dominated by its slowest sample.
This is why "the model beats the human baseline" and "the model is far below human performance" are frequently both true statements about the same model and the same task family. They are readings at different points on the surface. Before you argue about the gap, make both sides name their four coordinates — the disagreement usually dissolves into two different questions, one about capability and one about deployment.
The grader is not neutral, and it is biased in both directions at once.
Assume you fixed the first three coordinates. The grader still tilts the comparison, and the interesting part is that the two common graders tilt it opposite ways, which is why the literature can support any conclusion you want.
# Same work, two graders, two directions of bias automated checker (tests, exact match, SQL result diff) favours the worker optimising for the check costs the human, who solves the task as stated and loses points on an unstated convention LLM judge (rubric, pairwise, reference-free) favours structured, verbose, confident prose costs the human, whose terse correct answer reads as low-effort against the rubric # The only stable fix one grader, blind to authorship, applied to both, with the human transcripts in the SAME format
The blinding requirement is stronger than it sounds. Human submissions usually arrive as a pull request with a commit message; agent submissions arrive as a diff plus a transcript. A judge can tell which is which from formatting alone, and judge preference for agent-shaped output is a documented failure mode rather than a suspicion — this is the same calibration problem as judge calibration and meta-evaluation, with the extra wrinkle that here the two populations are systematically distinguishable.
Automated checkers have the complementary problem. A test suite encodes conventions the task statement did not, and the agent learns them from the repository while the human annotator reads the ticket. That is a real difference in performance on the deployed task, so it is not automatically a bias to be removed — but it must be reported, because "the agent passed the hidden tests more often" and "the agent solved the problem better" are different claims and only the first one was measured. Related: eval integrity and scorer gaming.
The one baseline the field actually built, and what it does and does not say.
The strongest human baselining effort in public agent evaluation is METR's time-horizon work, and it is worth understanding precisely because its method is the opposite of the usual one. Rather than asking "what fraction of tasks does the model pass", it estimates human expert completion time for each task, fits a logistic curve of model success against the logarithm of that human time, and reports the point where the curve crosses 50% — the 50% task-completion time horizon, expressed in hours of equivalent human effort.
- What it buys. The human measurement becomes the x-axis rather than a single comparison point, so the result is a statement about task difficulty calibrated in human hours, not a percentage that depends on the mix of easy and hard tasks in your set. Two benchmarks with different mixes produce comparable horizons; two benchmarks with different mixes do not produce comparable pass rates.
- What it cost. Real baselining runs, with multiple annotators per task and geometric-mean aggregation over successful attempts. The January 2026 Time Horizon 1.1 update expanded the long-horizon suite from 14 to 31 tasks, each estimated at eight or more human hours — which tells you the price of extending a calibrated baseline into the region everyone most wants measured.
- What it does not say. METR has published its own limitations note on the metric, and the honest reading is narrow: a 50% horizon is a statement about a distribution of software tasks with clean success criteria, not a claim that the model can be trusted for that long on your work. Reliability at the horizon is 50% by construction — see task horizon for why the number people want is the 80% or 95% horizon, which is much shorter.
Set against that, notice what the benchmark most quoted in agent procurement does not have. SWE-bench Verified has a human-validated task set — annotators confirmed the problems are solvable and the tests fair — which is a different thing from a human performance baseline on those same tasks under a time cap. So the widely repeated framing "model X now solves 70% of real GitHub issues, approaching expert level" contains a comparison that was never run. The saturation argument in benchmark landscape (2026) applies on top: once five models sit within a point of each other, the missing denominator is the only thing left that could make the number decision-relevant.
Build your own baseline for about two person-days.
The reason teams skip this is a belief that baselining is a research programme. On your own task set it is not, because you do not need a publishable estimate — you need a denominator stable enough to divide by twice a year. Twenty tasks and two annotators is enough to move you from no baseline to a defensible one.
# Minimum viable baseline, ~2 person-days 1 sample 20 tasks from the same pool the agent runs on stratify by your own difficulty proxy, not by outcome 2 fix the budget BEFORE anyone starts e.g. 45 min/task, hard stop, partial credit allowed 3 two annotators, disjoint halves, 4 tasks overlapping the overlap is your inter-annotator agreement 4 submissions normalised to the agent's output format strip commit messages, author names, styling 5 grade blind, with the agent's grader, in one sitting 6 publish: n, budget, tool policy, agreement, CI # Refresh cadence re-baseline when the task pool changes, not when the model changes. The baseline is a property of the tasks.
- Stratify before you see outcomes. Sampling tasks after you know which ones the agent fails produces a baseline that answers a question you already knew the answer to. The stratification variable should be something you can compute from the task, not from a run.
- Allow partial credit on both sides or neither. Mixed regimes are the most common quiet error: agents get pass/fail from a checker while humans get judged generously by a colleague who can see they were nearly there.
- Spend the overlap budget. Four shared tasks cost almost nothing and give you the one number that tells you whether your rubric is a rubric or a mood. If the two annotators disagree on the shared tasks, your agent score has the same instability and you have been reading noise — the machinery for that is in eval variance and statistical power.
- Pay for it once, amortise it for a year. A twenty-task baseline is a fixed cost against every model evaluation you run afterwards, which makes it cheap by the standards of cost of evaluation — and far cheaper than the alternative, which is a procurement decision made on a ratio with no denominator.
Report the ratio, not the score.
Once you have a baseline, change what you publish internally. A raw percentage invites comparison with other teams' raw percentages, which are not comparable; a ratio against a stated baseline forces the coordinates into the sentence and makes the claim auditable.
- Lead with the tuple. "0.90 of our internal baseline (2h cap, non-expert-in-this-repo annotators, same grader, n=6)" is longer than "37/60" and it is the only form of the sentence a reader can act on.
- Never compare across caps. An uncapped model score against a capped human baseline is the single most common way a ratio gets manufactured. If you cannot cap the agent the same way, report both numbers and refuse to divide.
- Keep the human transcripts. They are the most valuable artefact the exercise produces: a labelled set of correct trajectories on your own tasks, which feeds eval set maintenance and gives your judge something to calibrate against.
- Expect the ratio to be jagged. A single ratio over a mixed task pool hides the shape that matters — above baseline on some task families, far below on others, for reasons that do not track difficulty as a human would rank it. That is the jagged frontier, and the per-family breakdown is where the deployment decision actually gets made.
Do this today: take the headline number your team quotes most often and try to write its four coordinates. In most teams one is unknown (the grader was different for the two sides), one is wrong (the human figure came from a paper on a different task set), and two were never recorded. That exercise costs twenty minutes and it usually kills one slide. Then run the twenty-task baseline in STEP 5, and read reading agent benchmarks for how to apply the same scepticism to numbers you did not collect.