ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ziru, Chen, Shijie, Ning, Yuting, Zhang, Qianheng, Wang, Boshi, Yu, Botao, Li, Yifei, Liao, Zeyi, Wei, Chen, Lu, Zitong, Dey, Vishal, Xue, Mingyi, Baker, Frazier N., Burns, Benjamin, Adu-Ampratwum, Daniel, Huang, Xuhui, Ning, Xia, Gao, Song, Su, Yu, Sun, Huan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913767519944704
author Chen, Ziru
Chen, Shijie
Ning, Yuting
Zhang, Qianheng
Wang, Boshi
Yu, Botao
Li, Yifei
Liao, Zeyi
Wei, Chen
Lu, Zitong
Dey, Vishal
Xue, Mingyi
Baker, Frazier N.
Burns, Benjamin
Adu-Ampratwum, Daniel
Huang, Xuhui
Ning, Xia
Gao, Song
Su, Yu
Sun, Huan
author_facet Chen, Ziru
Chen, Shijie
Ning, Yuting
Zhang, Qianheng
Wang, Boshi
Yu, Botao
Li, Yifei
Liao, Zeyi
Wei, Chen
Lu, Zitong
Dey, Vishal
Xue, Mingyi
Baker, Frazier N.
Burns, Benjamin
Adu-Ampratwum, Daniel
Huang, Xuhui
Ning, Xia
Gao, Song
Su, Yu
Sun, Huan
contents The advancements of large language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about their true capabilities. In this work, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for evaluating language agents for data-driven scientific discovery. To ensure the scientific authenticity and real-world relevance of our benchmark, we extract 102 tasks from 44 peer-reviewed publications in four disciplines and engage nine subject matter experts to validate them. We unify the target output for every task to a self-contained Python program file and employ an array of evaluation metrics to examine the generated programs, execution results, and costs. Each task goes through multiple rounds of manual validation by annotators and subject matter experts to ensure its annotation quality and scientific plausibility. We also propose two effective strategies to mitigate data contamination concerns. Using ScienceAgentBench, we evaluate five open-weight and proprietary LLMs, each with three frameworks: direct prompting, OpenHands CodeAct, and self-debug. Given three attempts for each task, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge. In addition, we evaluate OpenAI o1-preview with direct prompting and self-debug, which can boost the performance to 42.2%, demonstrating the effectiveness of increasing inference-time compute but with more than 10 times the cost of other LLMs. Still, our results underscore the limitations of current language agents in generating code for data-driven discovery, let alone end-to-end automation for scientific research.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05080
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Chen, Ziru
Chen, Shijie
Ning, Yuting
Zhang, Qianheng
Wang, Boshi
Yu, Botao
Li, Yifei
Liao, Zeyi
Wei, Chen
Lu, Zitong
Dey, Vishal
Xue, Mingyi
Baker, Frazier N.
Burns, Benjamin
Adu-Ampratwum, Daniel
Huang, Xuhui
Ning, Xia
Gao, Song
Su, Yu
Sun, Huan
Computation and Language
Artificial Intelligence
Machine Learning
The advancements of large language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about their true capabilities. In this work, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for evaluating language agents for data-driven scientific discovery. To ensure the scientific authenticity and real-world relevance of our benchmark, we extract 102 tasks from 44 peer-reviewed publications in four disciplines and engage nine subject matter experts to validate them. We unify the target output for every task to a self-contained Python program file and employ an array of evaluation metrics to examine the generated programs, execution results, and costs. Each task goes through multiple rounds of manual validation by annotators and subject matter experts to ensure its annotation quality and scientific plausibility. We also propose two effective strategies to mitigate data contamination concerns. Using ScienceAgentBench, we evaluate five open-weight and proprietary LLMs, each with three frameworks: direct prompting, OpenHands CodeAct, and self-debug. Given three attempts for each task, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge. In addition, we evaluate OpenAI o1-preview with direct prompting and self-debug, which can boost the performance to 42.2%, demonstrating the effectiveness of increasing inference-time compute but with more than 10 times the cost of other LLMs. Still, our results underscore the limitations of current language agents in generating code for data-driven discovery, let alone end-to-end automation for scientific research.
title ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.05080