Young-Jun Lee, Seungone Kim, Jinheon Baek, Soojung Yang, Yoonho Lee, Dongyeop Kang
<aside> 🌟
Through this work, we aim to answer the question:
Can AI agents evaluate what constitutes good science as well as scientists can?
</aside>
Until last year, benchmarks based on competition problems, such as IMO, focused on how well reasoning models could solve given problems through deep reasoning, and performance on many of these benchmarks had largely saturated. As the agentic AI market has grown, a large number of agent benchmarks have emerged, including long-horizon agent benchmarks such as GAIA and tau3-bench; agentic coding benchmarks such as SWE-Bench Verified, TerminalBench 2.1, DeepSWE, and FrontierCode; and agentic computer-use benchmarks such as OSWorld 2.0. These recent benchmarks all evaluate agent capabilities in a static way. Some use deterministically verifiable scores as their headline metrics, while others use fixed rubrics created by domain experts and score outcomes through an LLM-as-a-Judge approach.
More recently, the rise of AI for Science has led to many proposed science benchmarks, but these also continue to adopt static evaluation methods: a single scalar headline metric based on a verifiable function or a scientist-verified rubric. Examples include PaperBench and NatureBench for paper reproduction; MLE-Bench for AI/ML engineering; FrontierScience for proposing follow-up research directions; and LifeSciBench, TerminalBench-Science, and FrontierPhysics for expert-curated scientific problems. Nevertheless, we believe that conventional static, fixed-target evaluation has limitations when assessing whether AI agents are doing “good science” in the way domain scientists would judge it.
Many scientific problems are open-ended: there is no reference answer, so no one knows in advance what a novel discovery will look like.
Therefore, as experiments progress, rubrics should evolve to reflect newly identified weaknesses, and AI agents should possess verification capabilities comparable to those of domain scientists. We believe this could be a starting point for the next stage of progress.
<aside>
🌟🌟🌟 All intellectual property rights to the outcomes generated by the agent for a proposed task will belong to the task proposer and reviewers. We will not claim ownership of these outcomes. You may use them to submit a paper to a Nature-family journal, and you are not required to include members of the co-lead team as co-authors.
</aside>
Guidelines for Task Collection: [Guideline] Track 4: SciVeri-Bench Task Proposal
One Interface for Collaboration. To create open-ended and novel scientific tasks, we consider the following points: