ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yujie, Yang, Zonglin, Xie, Tong, Ni, Jinjie, Gao, Ben, Li, Yuqiang, Tang, Shixiang, Ouyang, Wanli, Cambria, Erik, Zhou, Dongzhan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908978790793216
author Liu, Yujie
Yang, Zonglin
Xie, Tong
Ni, Jinjie
Gao, Ben
Li, Yuqiang
Tang, Shixiang
Ouyang, Wanli
Cambria, Erik
Zhou, Dongzhan
author_facet Liu, Yujie
Yang, Zonglin
Xie, Tong
Ni, Jinjie
Gao, Ben
Li, Yuqiang
Tang, Shixiang
Ouyang, Wanli
Cambria, Erik
Zhou, Dongzhan
contents Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. To address this gap, we introduce the first large-scale benchmark for evaluating LLMs on a sufficient set of scientific discovery sub-tasks-inspiration retrieval, hypothesis composition, and hypothesis ranking-where sufficient means that perfectly solving these sub-tasks perfectly solves the overall discovery task. We develop an automated LLM-based framework that extracts critical components-research questions, background surveys, inspirations, and hypotheses-from papers across 12 disciplines, with expert validation confirming its accuracy. To prevent data contamination, we focus exclusively on publications from 2024 onward, ensuring minimal overlap with LLM pretraining data; our automated framework further enables automatic extraction of even more recent papers as LLM pretraining cutoffs advance, supporting scalable and contamination-free automatic renewal of this discovery benchmark. Our evaluation shows that, across disciplines, LLMs excel at inspiration retrieval-an out-of-distribution task-suggesting their ability to surface novel knowledge associations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
Liu, Yujie
Yang, Zonglin
Xie, Tong
Ni, Jinjie
Gao, Ben
Li, Yuqiang
Tang, Shixiang
Ouyang, Wanli
Cambria, Erik
Zhou, Dongzhan
Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. To address this gap, we introduce the first large-scale benchmark for evaluating LLMs on a sufficient set of scientific discovery sub-tasks-inspiration retrieval, hypothesis composition, and hypothesis ranking-where sufficient means that perfectly solving these sub-tasks perfectly solves the overall discovery task. We develop an automated LLM-based framework that extracts critical components-research questions, background surveys, inspirations, and hypotheses-from papers across 12 disciplines, with expert validation confirming its accuracy. To prevent data contamination, we focus exclusively on publications from 2024 onward, ensuring minimal overlap with LLM pretraining data; our automated framework further enables automatic extraction of even more recent papers as LLM pretraining cutoffs advance, supporting scalable and contamination-free automatic renewal of this discovery benchmark. Our evaluation shows that, across disciplines, LLMs excel at inspiration retrieval-an out-of-distribution task-suggesting their ability to surface novel knowledge associations.
title ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
topic Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2503.21248