Young-Jun Lee, Seungone Kim, Jinheon Baek, Soojung Yang, Yoonho Lee, Dongyeop Kang
<aside> 🌟
Through this work, we aim to answer the question:
Can AI agents evaluate what constitutes good science as well as scientists can?
</aside>
Until last year, benchmarks based on competition problems, such as IMO, focused on how well reasoning models could solve given problems through deep reasoning, and performance on many of these benchmarks had largely saturated. As the agentic AI market has grown, a large number of agent benchmarks have emerged, including long-horizon agent benchmarks such as GAIA and tau3-bench; agentic coding benchmarks such as SWE-Bench Verified, TerminalBench 2.1, DeepSWE, and FrontierCode; and agentic computer-use benchmarks such as OSWorld 2.0. These recent benchmarks all evaluate agent capabilities in a static way. Some use deterministically verifiable scores as their headline metrics, while others use fixed rubrics created by domain experts and score outcomes through an LLM-as-a-Judge approach.
More recently, the rise of AI for Science has led to many proposed science benchmarks, but these also continue to adopt static evaluation methods: a single scalar headline metric based on a verifiable function or a scientist-verified rubric. Examples include PaperBench and NatureBench for paper reproduction; MLE-Bench for AI/ML engineering; FrontierScience for proposing follow-up research directions; and LifeSciBench, TerminalBench-Science, and FrontierPhysics for expert-curated scientific problems.
However, static evaluation metrics alone are not sufficient to verify whether an AI agent’s output constitutes good science. Many scientific problems, such as causal mechanism discovery, are open-ended: a satisfactory solution has yet to be discovered, and there is no reference answer that defines what constitutes good science [1,2]. For such problems, no one knows what a novel discovery or good science will look like before experiments are conducted. Therefore, static and fixed evaluation metrics, even with rubrics curated by scientists, are not sufficient to verify good science.
Over time, as experiments are conducted and results accumulate, domain scientists (1) identify weaknesses or gaps in the experimental results, (2) use these observations to progressively refine their criteria for what constitutes good science in the context of those experiments, and (3) judge which results are better based on their experience and intuition (e.g., whether earlier results were better or which of two experiments produced better results).
However, most existing science benchmarks have focused on agents’ problem-solving capabilities rather than their verification capabilities. Although it is important to evaluate how well AI agents can solve scientific problems compared with scientists, it is equally important to evaluate how well they can match scientists in verifying whether an outcome constitutes good science or a novel discovery.
<aside>
🌟🌟🌟 All intellectual property rights to the outcomes generated by the agent for a proposed task will belong to the task proposer and reviewers. We will not claim ownership of these outcomes. You may use them to submit a paper to a Nature-family journal, and you are not required to include members of the co-lead team as co-authors.
</aside>
Guidelines for Task Collection: [Guideline] Track 4: SciVeri-Bench Task Proposal
One Interface for Collaboration. To create open-ended and novel scientific tasks, we consider the following points: