HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Haokun, Huang, Sicong, Hu, Jingyu, Zhou, Yangqiaoyu, Tan, Chenhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914320306143232
author Liu, Haokun
Huang, Sicong
Hu, Jingyu
Zhou, Yangqiaoyu
Tan, Chenhao
author_facet Liu, Haokun
Huang, Sicong
Hu, Jingyu
Zhou, Yangqiaoyu
Tan, Chenhao
contents There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematically evaluate methods for hypothesis generation? To address this, we introduce HypoBench, a novel benchmark designed to evaluate LLMs and hypothesis generation methods across multiple aspects, including practical utility, generalizability, and hypothesis discovery rate. HypoBench includes 7 real-world tasks and 5 synthetic tasks with 194 distinct datasets. We evaluate four state-of-the-art LLMs combined with six existing hypothesis-generation methods. Overall, our results suggest that existing methods are capable of discovering valid and novel patterns in the data. However, the results from synthetic datasets indicate that there is still significant room for improvement, as current hypothesis generation methods do not fully uncover all relevant or meaningful patterns. Specifically, in synthetic settings, as task difficulty increases, performance significantly drops, with best models and methods only recovering 38.8% of the ground-truth hypotheses. These findings highlight challenges in hypothesis generation and demonstrate that HypoBench serves as a valuable resource for improving AI systems designed to assist scientific discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11524
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
Liu, Haokun
Huang, Sicong
Hu, Jingyu
Zhou, Yangqiaoyu
Tan, Chenhao
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematically evaluate methods for hypothesis generation? To address this, we introduce HypoBench, a novel benchmark designed to evaluate LLMs and hypothesis generation methods across multiple aspects, including practical utility, generalizability, and hypothesis discovery rate. HypoBench includes 7 real-world tasks and 5 synthetic tasks with 194 distinct datasets. We evaluate four state-of-the-art LLMs combined with six existing hypothesis-generation methods. Overall, our results suggest that existing methods are capable of discovering valid and novel patterns in the data. However, the results from synthetic datasets indicate that there is still significant room for improvement, as current hypothesis generation methods do not fully uncover all relevant or meaningful patterns. Specifically, in synthetic settings, as task difficulty increases, performance significantly drops, with best models and methods only recovering 38.8% of the ground-truth hypotheses. These findings highlight challenges in hypothesis generation and demonstrate that HypoBench serves as a valuable resource for improving AI systems designed to assist scientific discovery.
title HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
topic Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2504.11524