Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koneru, Sai, Wu, Jian, Rajtmajer, Sarah
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911812636639232
author Koneru, Sai
Wu, Jian
Rajtmajer, Sarah
author_facet Koneru, Sai
Wu, Jian
Rajtmajer, Sarah
contents Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies in the social sciences. We compare the performance of LLMs to several state-of-the-art benchmarks and highlight opportunities for future research in this area. The dataset is available at https://github.com/Sai90000/ScientificHypothesisEvidencing.git
format Preprint
id arxiv_https___arxiv_org_abs_2309_06578
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
Koneru, Sai
Wu, Jian
Rajtmajer, Sarah
Computation and Language
Artificial Intelligence
Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies in the social sciences. We compare the performance of LLMs to several state-of-the-art benchmarks and highlight opportunities for future research in this area. The dataset is available at https://github.com/Sai90000/ScientificHypothesisEvidencing.git
title Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.06578