Young-Jun Lee, Seungone Kim, Jinheon Baek, Soojung Yang, Yoonho Lee, Dongyeop Kang

<aside> 🌟

Through this work, we aim to answer the question:

Can AI agents evaluate what constitutes good science as well as scientists can?

</aside>

Current agent benchmarks rely on static evaluation

Until last year, benchmarks based on competition problems, such as IMO, focused on how well reasoning models could solve given problems through deep reasoning, and performance on many of these benchmarks had largely saturated. As the agentic AI market has grown, a large number of agent benchmarks have emerged, including long-horizon agent benchmarks such as GAIA and tau3-bench; agentic coding benchmarks such as SWE-Bench Verified, TerminalBench 2.1, DeepSWE, and FrontierCode; and agentic computer-use benchmarks such as OSWorld 2.0. These recent benchmarks all evaluate agent capabilities in a static way. Some use deterministically verifiable scores as their headline metrics, while others use fixed rubrics created by domain experts and score outcomes through an LLM-as-a-Judge approach.

Science is alive, yet its evaluation remains static.

More recently, the rise of AI for Science has led to many proposed science benchmarks, but these also continue to adopt static evaluation methods: a single scalar headline metric based on a verifiable function or a scientist-verified rubric. Examples include PaperBench and NatureBench for paper reproduction; MLE-Bench for AI/ML engineering; FrontierScience for proposing follow-up research directions; and LifeSciBench, TerminalBench-Science, and FrontierPhysics for expert-curated scientific problems. Nevertheless, we believe that conventional static, fixed-target evaluation has limitations when assessing whether AI agents are doing “good science” in the way domain scientists would judge it.

Static evaluation metrics are not suited to verifying “good science”

Many scientific problems are open-ended: there is no reference answer, so no one knows in advance what a novel discovery will look like.

  1. Rubrics are not fully articulated until scientific experiments are conducted. Scientists do not necessarily create effective rubrics from the outset. Before running an experiment, it is difficult to judge which findings are truly valuable and which directions are unpromising. It is therefore inappropriate to keep using a fixed rubric established when the task was proposed—that is, before the experiment was run.
  2. AI agents do not know where the results fall short. Because AI agents do not know which aspects of experimental results are lacking, it is difficult for them to determine how to improve the experiments or generate rubrics. This is especially challenging for open-ended tasks, where the correct answer is unknown.
  3. AI agents should be able to judge which experimental results are good: Scientists draw on their experience and intuition to make these judgments, but AI agents find them difficult. Such judgment is necessary for refining the design of subsequent experiments and pursuing novel discoveries.

Therefore, as experiments progress, rubrics should evolve to reflect newly identified weaknesses, and AI agents should possess verification capabilities comparable to those of domain scientists. We believe this could be a starting point for the next stage of progress.

Contributions

<aside>

🌟🌟🌟 All intellectual property rights to the outcomes generated by the agent for a proposed task will belong to the task proposer and reviewers. We will not claim ownership of these outcomes. You may use them to submit a paper to a Nature-family journal, and you are not required to include members of the co-lead team as co-authors.

</aside>

SciVeri-Bench: Benchmarking AI Agents for Scientific Verification

Guidelines for Task Collection: [Guideline] Track 4: SciVeri-Bench Task Proposal

Benchmark Design Principle

One Interface for Collaboration. To create open-ended and novel scientific tasks, we consider the following points: