CiteME: Can Language Models Accurately Cite Scientific Claims?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Press, Ori, Hochlehnert, Andreas, Prabhu, Ameya, Udandarao, Vishaal, Press, Ofir, Bethge, Matthias
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929573430558720
author Press, Ori
Hochlehnert, Andreas
Prabhu, Ameya
Udandarao, Vishaal
Press, Ofir
Bethge, Matthias
author_facet Press, Ori
Hochlehnert, Andreas
Prabhu, Ameya
Udandarao, Vishaal
Press, Ofir
Bethge, Matthias
contents Thousands of new scientific papers are published each month. Such information overload complicates researcher efforts to stay current with the state-of-the-art as well as to verify and correctly attribute claims. We pose the following research question: Given a text excerpt referencing a paper, could an LM act as a research assistant to correctly identify the referenced paper? We advance efforts to answer this question by building a benchmark that evaluates the abilities of LMs in citation attribution. Our benchmark, CiteME, consists of text excerpts from recent machine learning papers, each referencing a single other paper. CiteME use reveals a large gap between frontier LMs and human performance, with LMs achieving only 4.2-18.5% accuracy and humans 69.7%. We close this gap by introducing CiteAgent, an autonomous system built on the GPT-4o LM that can also search and read papers, which achieves an accuracy of 35.3\% on CiteME. Overall, CiteME serves as a challenging testbed for open-ended claim attribution, driving the research community towards a future where any claim made by an LM can be automatically verified and discarded if found to be incorrect.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CiteME: Can Language Models Accurately Cite Scientific Claims?
Press, Ori
Hochlehnert, Andreas
Prabhu, Ameya
Udandarao, Vishaal
Press, Ofir
Bethge, Matthias
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Thousands of new scientific papers are published each month. Such information overload complicates researcher efforts to stay current with the state-of-the-art as well as to verify and correctly attribute claims. We pose the following research question: Given a text excerpt referencing a paper, could an LM act as a research assistant to correctly identify the referenced paper? We advance efforts to answer this question by building a benchmark that evaluates the abilities of LMs in citation attribution. Our benchmark, CiteME, consists of text excerpts from recent machine learning papers, each referencing a single other paper. CiteME use reveals a large gap between frontier LMs and human performance, with LMs achieving only 4.2-18.5% accuracy and humans 69.7%. We close this gap by introducing CiteAgent, an autonomous system built on the GPT-4o LM that can also search and read papers, which achieves an accuracy of 35.3\% on CiteME. Overall, CiteME serves as a challenging testbed for open-ended claim attribution, driving the research community towards a future where any claim made by an LM can be automatically verified and discarded if found to be incorrect.
title CiteME: Can Language Models Accurately Cite Scientific Claims?
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2407.12861