Scientific-discovery agents.
Two benchmarks posted to arXiv in the opening days of October 2026 settle the design question before you write any code: on EurekaBench an agent reached 47.4% predictive accuracy against a human scientist's 48.8%, and scored 29.4% on the insight the task was built around against the humans' 69.7%. Build this agent to maximise the metric it can already win and you ship a machine that produces unexplained correlations at expert accuracy — the one output a scientist cannot use and cannot publish. The explanation is the product, so make it the scored artifact, and make the amount of methodological guidance you supplied an explicit logged parameter rather than an accident of the prompt.
Three jobs that look like one, and must not be one agent.
"An agent for science" bundles three tasks with different success criteria, different verification, and radically different difficulty. Bundling them is the characteristic first mistake, because it means every failure is unattributable: you cannot tell whether the agent got the physics wrong or the file format wrong.
- Execute a known method. Prepare inputs for a simulation or instrument whose procedure is documented, run it, collect outputs. Verification is exact — the inputs either match a reference or do not. This is largely a tool-driving problem and it is the part that works today.
- Analyse given outputs. Take results that already exist and extract the quantity the study was after. Verification is still mechanical if you have reproduced ground truth. This is where most of the usable value currently sits.
- Propose a mechanism. Look at observations and produce an explanation that generalises. Verification is a judgement about whether the explanation supports the insights the phenomenon actually has — and this is the job where agents are furthest behind, by a factor the benchmark numbers in the lede make concrete.
Ship the first two as separate agents with separate evals before attempting the third. The ordering is not conservatism; it is that the third job's failures are only legible once the first two are known-good, and a bundled agent hides exactly that.
Pre-run the expensive thing; grade with rules, not a judge.
The instinct is to give the agent a compute allocation and let it run simulations. That makes every eval cost real money and real hours, which means you run it rarely, which means you never find out whether a change helped. CompMat-Bench — 94 tasks drawn from recently published computational-materials studies — takes the other route: the expensive simulations are pre-run, the reproduced inputs and results become the ground truth, and grading uses fixed rules rather than an LLM judge. Copy that shape for your own harness.
# Two harness shapes for the same agent LIVE COMPUTE agent submits jobs, waits, reads results eval cost hours + real cluster spend per run grading whatever the agent says it found consequence you run the eval monthly and fly blind between PRE-RUN GROUND TRUTH outputs already computed and stored eval cost minutes, repeatable, parallel grading fixed rules against reproduced values consequence you can gate a prompt change on it # What this buys beyond speed a rule-graded score cannot be talked into agreeing with the agent (an LLM judge reading the agent's own narrative can, and will)
Keep live compute for the small set of end-to-end runs that have to be real, and treat it as a separately budgeted capability with its own approval — a job submission is an expensive, slow, externally visible side effect, which puts it in the same class as any other irreversible tool call. The judge-free grading matters for a second reason: an agent that writes its own narrative of what it found is the generator and the persuader, and an LLM judge reading that narrative scores the writing. See the generator–verifier gap and LLM-as-judge for agents for when a judge is defensible at all.
Make guidance a parameter you set, not a habit you drift into.
This is the move that distinguishes a measurable discovery agent from a demo. CompMat-Bench evaluates under four conditions — single tasks and multi-step workflows, each with full or reduced methodological guidance — and pass rates that sit at 66.0–90.4% on single tasks with full guidance fall away as the workflow lengthens and the guidance is withdrawn. Most of the remaining failures are scientific-reasoning errors rather than software errors, which is the finding that should change your build: a better harness does not close that gap.
So define guidance levels explicitly, store the level on every run, and never let the prompt drift between them silently.
# Guidance levels, declared per task and logged per run L0 goal only "explain the anomaly in this dataset" L1 + method family "use DFT"; no parameters L2 + parameters functional, cutoffs, convergence criteria L3 + full procedure the published methods section, verbatim # The rule that keeps the number honest a result reported without its guidance level is not a result L3 pass rate is a statement about your prompt L0 pass rate is a statement about the agent
Two practical consequences. In production, run at the highest guidance level that is honest — you want the work done, and withheld guidance is just a self-imposed handicap. In evaluation, run at several levels and report the curve, because a single number at L3 tells you how good your methods section is. The same distinction shows up in human baselines in agent evals: the comparison is only meaningful when both sides got the same help.
Score the explanation separately from the prediction.
EurekaBench's design makes the split concrete: 26 expert-verified long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics and plasma physics, carrying 306 scientific insights that a correct mechanism should support. The headline result is the decoupling. Predictive accuracy at human parity, insight at well under half the human level, in the same runs. An agent optimising a loss on held-out observations will walk straight into that corner, because fitting is the thing gradient-shaped effort is good at.
Build two scores and keep them separate on the dashboard.
- Predictive score — does it forecast held-out observations? Mechanical, cheap, and the one your agent will optimise if you let it be the only number.
- Insight score — does the proposed mechanism entail the specific claims a domain expert says the phenomenon supports? Enumerate those claims in advance, as EurekaBench does, so the scoring is a checklist against a pre-registered list rather than an impression of the write-up.
- Falsifiability — does the explanation make a prediction that would fail if it were wrong? An unfalsifiable mechanism that fits perfectly is the characteristic failure output, and it is easy to detect precisely because it forbids nothing.
Pre-registering the insight list is the cheapest high-leverage thing on this page. It costs a domain expert an afternoon per phenomenon and it converts "the agent wrote a plausible explanation" into a count. Without it you will grade narratives, and narratives are what an LLM is best at producing and worst at being scored on.
Provenance, or the result is not a result.
A scientific claim is worth exactly what its audit trail is worth, and the agent is the part of the pipeline with no memory of its own. Every number the agent reports has to be reachable backwards to the thing that produced it, automatically, without the agent being asked to remember.
- Inputs by content hash. Every dataset, input file and parameter set referenced by digest, not by filename — filenames get overwritten, and a result attributed to
run_final_v2.datis attributable to nothing. Same discipline as pinning and verification. - Code and environment version per step. The solver version, the library versions, the random seed. Scientific software changes defaults between minor releases, and a result that cannot name its solver version cannot be reproduced or defended.
- The analysis as code, not as chat. The agent should emit a script that produces the figure, and the figure should be regenerated by running the script in your harness — not pasted from the transcript. See notebook and data-science agents, where the same gap between "what the agent reported" and "what reruns" is the whole subject.
- Nondeterminism declared, not discovered. Sampling, parallel reductions and GPU nondeterminism all move the last digits; state the tolerance your comparison uses, as in reproducibility and determinism.
Hold the line that a claim without this trail does not leave the system, even when it is right. The cost of relaxing it is not a wrong paper; it is that nobody can tell your right results from your wrong ones afterwards.
Where the human sits, and what you can actually ship this year.
Put the human at the two boundaries where the agent is weakest and the cost of being wrong is highest: choosing what question is worth asking, and deciding whether an explanation is an explanation. In between — preparing inputs, driving tools, extracting quantities, producing the figure, writing the first draft of the methods — the agent earns its place.
# A shippable scope, by job (Step 1) and guidance level (Step 3) execute known method L2–L3 autonomous, rule-graded, high volume analyse given outputs L1–L2 autonomous with provenance gate propose a mechanism L0–L1 drafts only, expert scores the insight # The metric to put on the wall not "discoveries made" expert-accepted insights per expert review-hour (and the denominator is the one that decides whether this is worth it)
That denominator is the honest unit for this whole class of agent. A system that generates twelve plausible mechanisms an hour and costs a scientist a day to triage has moved work, not done it — the same arithmetic that governs the cost of human review everywhere else. Measure the review hours from the first week, before the demo convinces anyone that throughput is the goal.
Start here, in this order: pick one published study in your domain, reproduce its inputs and results and freeze them as ground truth, write the rule-based grader, then run your agent at L3 and at L0 on the same task. The spread between those two numbers is the most informative measurement you will take all quarter — it tells you how much of your current demo is the agent and how much is you. If the L0 number is near zero, you have an execution agent, which is a genuinely useful product; just do not ship it as a discovery one.
Related: research agents for literature-facing work rather than experiments, trajectory and process evaluation for grading the path rather than the answer, and task horizon for why the multi-step condition is where the pass rate goes.