AI Blog

It matched the scientist and missed the point

Two benchmarks posted to arXiv in the opening days of October 2026 turn the two things every other agent eval holds constant into variables — how much guidance the harness supplied, and whether the score rewards a prediction or an explanation. Both move the number by tens of points, and one of them reports an agent at 47.4% predictive accuracy against a human scientist’s 48.8% while scoring 29.4% against 69.7% on the insight the task was built around.

By Agentic AI Wiki 12 min read

An agent forecast held-out scientific observations about as well as the scientist who published them — 47.4% against 48.8% — and then scored 29.4% against that scientist's 69.7% on whether its explanation accounted for what was going on. Same runs, same tasks, forty points apart. That split, and a second benchmark that varies how much methodological help the harness hands over, together make a claim about every agent score you have ever read: it is a reading taken with the help level held at one setting and the scored object chosen for you. Change either and the number moves by tens of points, which means the headline figure in your own eval deck is answering a question you did not ask.

At a glance

Two papers from the opening days of October 2026, both aimed at scientific work rather than coding, both making a variable out of something benchmarks normally fix.

BenchmarkShapeWhat it turns into a variableThe number that matters
CompMat-Bench (arXiv 2610.00636) 94 tasks from recently published computational-materials studies How much methodological guidance the agent is given, crossed with task length 66.0–90.4% pass on single tasks with full guidance — falling as either axis moves
EurekaBench (arXiv 2610.00492) 26 expert-verified long-horizon tasks across six sciences, carrying 306 insights Whether the score rewards a prediction or the explanation behind it 47.4% predictive vs 48.8% human; 29.4% insight vs 69.7% human
CompMat-Bench evaluation conditions A two-by-two grid of evaluation conditions. Rows are single tasks and multi-step workflows; columns are full and reduced methodological guidance. The single-task, full-guidance cell is the only one carrying the published headline range of 66.0 to 90.4 percent pass rate; pass rates decline toward the workflow and reduced-guidance cells. 94 tasks, four conditions — the headline number lives in one cell FULL GUIDANCE REDUCED GUIDANCE SINGLE TASK WORKFLOW 66.0–90.4% published range, three LLMs lower method withheld lower steps compound lowest both withdrawn Most remaining failures are scientific-reasoning errors, not software errors — grading is rule-based, with no LLM judge.
The published range belongs to one of four cells. Three benchmarks in one, and only the easy corner gets quoted.

Your benchmark number is a guidance-level reading

CompMat-Bench's design is the interesting part, more than its scores. It draws 94 tasks from real published studies, pre-runs the expensive simulations so the reproduced inputs and results can serve as ground truth, and grades with fixed rules rather than an LLM judge. Then it evaluates under four conditions: single tasks and multi-step workflows, each with full or with reduced methodological guidance.

That second axis is the one nobody else varies. Every mainstream agent benchmark supplies a fixed amount of help — a task statement of a particular specificity, tool docs of a particular quality, an environment pre-configured to a particular degree — and then reports one number. The help is a constant, so it disappears from the result. CompMat-Bench makes it a treatment, and the treatment effect is large: pass rates of 66.0–90.4% on single tasks with full guidance decline as workflows lengthen and guidance is withdrawn.

Read that as a measurement property rather than a finding about materials science. If a benchmark's score is sensitive to how much procedural help it hands over, then two benchmarks that disagree about a model may simply be set at different help levels, and a leaderboard that ranks harnesses is partly ranking how much their authors wrote into the task statement. The honest unit is a curve over guidance, not a point — and essentially nobody reports one.

47.4 against 48.8, and 29.4 against 69.7

EurekaBench: predictive accuracy against insight score Grouped bars comparing an agent and a human reference on two axes. On predictive accuracy the agent scores 47.4 percent against the human 48.8 percent, near parity. On the scientific-insight score the agent scores 29.4 percent against the human 69.7 percent, a gap of more than forty points. Score (%), agent vs human reference 20 40 60 80 100 Predictive accuracy agent 47.4 human 48.8 Scientific-insight score agent 29.4 human 69.7 Same runs, same tasks — 40.3 points apart on the axis the task was built around.
Near parity on the axis that is easy to grade; a forty-point hole on the axis the task was built around.

EurekaBench scores the other thing. Its 26 tasks each present observations from a real study — neuroscience, computer science, chemistry, astrophysics, geophysics, plasma physics — and ask the agent to discover the mechanism that explains them. Crucially, the benchmark carries 306 scientific insights that a correct mechanism is expected to support, enumerated in advance by experts. So the explanation is not graded as prose. It is graded as a checklist against a pre-registered list of things a right answer has to entail.

GPT 6 Astra lands at 47.4% predictive accuracy against a human reference of 48.8% — close enough that, on the metric most agent evaluations would report, you would call it parity and move on. On the insight score it manages 29.4% where human-discovered mechanisms score 69.7%. The paper's framing is that agents over-fixate on optimising predictive accuracy, in places surpassing human scientists, while falling far short at deriving the insight.

The mechanism behind that is not mysterious and it is not specific to science. Predictive accuracy is a differentiable-ish target with abundant signal: the agent can try things, check against held-out data, and iterate. "Does this explanation entail the right consequences" offers almost no per-step signal, cannot be checked by running something, and rewards a kind of restraint — refusing to claim a mechanism the data does not pin down. Any effort that can be spent fitting will be spent fitting. Give an agent one scalar to maximise and it will find the cheapest route to it, which is exactly the reward-design problem we already know about from reward design and hacking, arriving this time at evaluation rather than training.

Why a better harness does not close this gap

The reflex when an agent benchmark comes out low is to blame the scaffolding: better tool docs, more retries, a planner, a verifier pass. Two details in these papers close that escape route.

First, CompMat-Bench reports that most remaining failures are scientific-reasoning errors rather than software mistakes. The agent is not failing to drive the simulator; it is choosing the wrong physics. Harness work moves software failures and leaves reasoning failures where they are, so the ceiling you hit by improving the harness is lower than the gap you are trying to close.

Second, the grading in CompMat-Bench is rule-based against reproduced ground truth, with no LLM judge, and EurekaBench's insight scoring runs against an expert-enumerated list. That removes the usual consolation — that the judge misunderstood a correct answer. A judge-free grader cannot be talked round by a confident write-up, which matters more than it sounds: an agent that produces both the result and the narrative about the result is the generator and the advocate at once, and an LLM judge reading that narrative is partly scoring the advocacy. We wrote about that asymmetry in the generator–verifier gap; these two benchmarks are what it looks like when you design it out.

There is a real limitation worth stating. Both papers evaluate a small number of agent configurations — three LLMs in CompMat-Bench's reported range — on domains where published ground truth exists, which biases toward work that was publishable and therefore tractable. Neither is a general statement about agent capability. The methodological point survives that limitation intact, because it is about what the measurement is sensitive to, not about where any particular model sits.

Four channels of help, and your eval holds three of them fixed

Where guidance enters an agent evaluation A diagram of four channels through which help reaches an agent during an evaluation: the task statement, the method specification, the tool and environment scaffolding, and the grader's tolerance. All four feed the agent loop, which produces a run that is scored. The diagram marks which channels a conventional benchmark holds fixed and which the two new benchmarks vary. FOUR CHANNELS OF HELP Task statement what counts as done usually fixed Method specification parameters, procedure, hints CompMat-Bench varies this Tools & environment pre-run results, retries, limits usually fixed What the grader rewards prediction, or explanation EurekaBench splits this Agent loop plan, call tools, read results, write Prediction checkable against held-out data Explanation checkable against listed insights A single reported score is a reading taken with all four channels held at one setting. Move the second channel and the number moves; split the fourth and one score becomes two that disagree.
Help arrives through four channels. A single reported score fixes all four at one setting and then forgets it did.

Laid out that way, the two papers are varying channels two and four of four. The other two — how the task is stated, and how much the environment was pre-arranged — remain constants in essentially all published agent evaluation, including these. That is worth knowing because it bounds what any single score can mean, and because channels two and four are the two you can vary cheaply in your own suite this week.

Three changes follow, in order of how much they cost you.

  • Log the guidance level on every run, and report at two of them. Define levels explicitly — goal only, goal plus method family, goal plus parameters, full procedure — and run your suite at the top and the bottom. The spread between those two numbers tells you how much of your result is the agent and how much is your prompt. It is the single most informative number most teams are not collecting.
  • Decide, in writing, which object you are scoring. Outcome or explanation. If both matter, score both and keep them on separate axes, because a single blended score lets a model buy the easy half. This is the outcome-versus-trajectory question with a sharper edge: the agent can be right and unable to say why, and whether that is a pass is a product decision you should make on purpose.
  • Enumerate the entailments in advance. EurekaBench's 306 insights are the expensive, boring, decisive part of its design. For your own domain it is an afternoon with a domain expert per scenario, and it converts "the agent gave a plausible explanation" into a count you can regress against. Without it you grade narratives, which is the thing a language model is best at producing and worst at being scored on.

When to pick which

If you are deciding…Read CompMat-Bench's designRead EurekaBench's design
Whether your agent can do a documented procedure reliablyYes — pre-run ground truth, rule-based grading, guidance held highNo — wrong question
Whether your demo is the agent or your promptYes — the four-condition cross is the whole methodPartly
Whether an answer you cannot explain is acceptableNoYes — this is exactly the split it measures
Whether to trust an LLM judge on your own evalsYes — it shows a rule-based alternative is reachableYes — pre-enumerated entailments instead of impressions
How a model ranks against other modelsNo — three configurations, one domainNo — 26 tasks

FAQ

Does an agent at 47.4% predictive accuracy actually beat human scientists?

Not on that number — 47.4% sits just under the 48.8% human reference, which is parity rather than a win. EurekaBench's broader framing is that agents over-optimise predictive accuracy and in places surpass human scientists on it; the load-bearing result is the decoupling between that axis and the insight axis, not a victory on either.

Isn't "withholding guidance" just making the benchmark artificially hard?

In production, no — you should give an agent all the help you honestly have. In evaluation it is the only way to separate the agent's contribution from yours. A pass rate measured with the full published methods section pasted into the prompt is a statement about that methods section.

Do these results say anything about coding agents?

The methodology does. The guidance axis exists in any agent eval — a SWE-style task with a reproduction script, a failing test and a hint is a different condition from an issue title alone, and both get reported as one score. The domain-reasoning finding does not transfer; that one is about materials physics.

Should I stop using an LLM judge?

No, but you should know what you lose. Where ground truth can be reproduced and compared by rule, a rule beats a judge — and these papers show how much of scientific evaluation is in that category. Where you keep a judge, calibrate it against a labelled set and re-check the agreement rate whenever the judge model changes.

What is the smallest version of this I can run?

Take ten tasks in your own suite, write a stripped task statement and a fully specified one for each, run both, and put the two pass rates side by side. That one table is the finding, in your domain, for a day's work.

Further reading

On this wiki:

Sources: