Large Language Models for Automated Open-domain Scientific Hypotheses Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zonglin, Du, Xinya, Li, Junxian, Zheng, Jie, Poria, Soujanya, Cambria, Erik
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929382353797120
author Yang, Zonglin
Du, Xinya
Li, Junxian
Zheng, Jie
Poria, Soujanya
Cambria, Erik
author_facet Yang, Zonglin
Du, Xinya
Li, Junxian
Zheng, Jie
Poria, Soujanya
Cambria, Erik
contents Hypothetical induction is recognized as the main reasoning type when scientists make observations about the world and try to propose hypotheses to explain those observations. Past research on hypothetical induction is under a constrained setting: (1) the observation annotations in the dataset are carefully manually handpicked sentences (resulting in a close-domain setting); and (2) the ground truth hypotheses are mostly commonsense knowledge, making the task less challenging. In this work, we tackle these problems by proposing the first dataset for social science academic hypotheses discovery, with the final goal to create systems that automatically generate valid, novel, and helpful scientific hypotheses, given only a pile of raw web corpus. Unlike previous settings, the new dataset requires (1) using open-domain data (raw web corpus) as observations; and (2) proposing hypotheses even new to humanity. A multi-module framework is developed for the task, including three different feedback mechanisms to boost performance, which exhibits superior performance in terms of both GPT-4 based and expert-based evaluation. To the best of our knowledge, this is the first work showing that LLMs are able to generate novel (''not existing in literature'') and valid (''reflecting reality'') scientific hypotheses.
format Preprint
id arxiv_https___arxiv_org_abs_2309_02726
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
Yang, Zonglin
Du, Xinya
Li, Junxian
Zheng, Jie
Poria, Soujanya
Cambria, Erik
Computation and Language
Artificial Intelligence
Hypothetical induction is recognized as the main reasoning type when scientists make observations about the world and try to propose hypotheses to explain those observations. Past research on hypothetical induction is under a constrained setting: (1) the observation annotations in the dataset are carefully manually handpicked sentences (resulting in a close-domain setting); and (2) the ground truth hypotheses are mostly commonsense knowledge, making the task less challenging. In this work, we tackle these problems by proposing the first dataset for social science academic hypotheses discovery, with the final goal to create systems that automatically generate valid, novel, and helpful scientific hypotheses, given only a pile of raw web corpus. Unlike previous settings, the new dataset requires (1) using open-domain data (raw web corpus) as observations; and (2) proposing hypotheses even new to humanity. A multi-module framework is developed for the task, including three different feedback mechanisms to boost performance, which exhibits superior performance in terms of both GPT-4 based and expert-based evaluation. To the best of our knowledge, this is the first work showing that LLMs are able to generate novel (''not existing in literature'') and valid (''reflecting reality'') scientific hypotheses.
title Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.02726