CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ou, Jiefu, Walden, William Gantt, Sanders, Kate, Jiang, Zhengping, Sun, Kaiser, Cheng, Jeffrey, Jurayj, William, Wanner, Miriam, Liang, Shaobo, Morgan, Candice, Han, Seunghoon, Wang, Weiqi, May, Chandler, Recknor, Hannah, Khashabi, Daniel, Van Durme, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909555353452544
author Ou, Jiefu
Walden, William Gantt
Sanders, Kate
Jiang, Zhengping
Sun, Kaiser
Cheng, Jeffrey
Jurayj, William
Wanner, Miriam
Liang, Shaobo
Morgan, Candice
Han, Seunghoon
Wang, Weiqi
May, Chandler
Recknor, Hannah
Khashabi, Daniel
Van Durme, Benjamin
author_facet Ou, Jiefu
Walden, William Gantt
Sanders, Kate
Jiang, Zhengping
Sun, Kaiser
Cheng, Jeffrey
Jurayj, William
Wanner, Miriam
Liang, Shaobo
Morgan, Candice
Han, Seunghoon
Wang, Weiqi
May, Chandler
Recknor, Hannah
Khashabi, Daniel
Van Durme, Benjamin
contents A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensuring that these reviews are sound and grounded in the papers' claims remains challenging. To facilitate LLM benchmarking on these challenges, we introduce CLAIMCHECK, an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews mined from OpenReview. CLAIMCHECK is richly annotated by ML experts for weakness statements in the reviews and the paper claims that they dispute, as well as fine-grained labels of the validity, objectivity, and type of the identified weaknesses. We benchmark several LLMs on three claim-centric tasks supported by CLAIMCHECK, requiring models to (1) associate weaknesses with the claims they dispute, (2) predict fine-grained labels for weaknesses and rewrite the weaknesses to enhance their specificity, and (3) verify a paper's claims with grounded reasoning. Our experiments reveal that cutting-edge LLMs, while capable of predicting weakness labels in (2), continue to underperform relative to human experts on all other tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
Ou, Jiefu
Walden, William Gantt
Sanders, Kate
Jiang, Zhengping
Sun, Kaiser
Cheng, Jeffrey
Jurayj, William
Wanner, Miriam
Liang, Shaobo
Morgan, Candice
Han, Seunghoon
Wang, Weiqi
May, Chandler
Recknor, Hannah
Khashabi, Daniel
Van Durme, Benjamin
Computation and Language
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensuring that these reviews are sound and grounded in the papers' claims remains challenging. To facilitate LLM benchmarking on these challenges, we introduce CLAIMCHECK, an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews mined from OpenReview. CLAIMCHECK is richly annotated by ML experts for weakness statements in the reviews and the paper claims that they dispute, as well as fine-grained labels of the validity, objectivity, and type of the identified weaknesses. We benchmark several LLMs on three claim-centric tasks supported by CLAIMCHECK, requiring models to (1) associate weaknesses with the claims they dispute, (2) predict fine-grained labels for weaknesses and rewrite the weaknesses to enhance their specificity, and (3) verify a paper's claims with grounded reasoning. Our experiments reveal that cutting-edge LLMs, while capable of predicting weakness labels in (2), continue to underperform relative to human experts on all other tasks.
title CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
topic Computation and Language
url https://arxiv.org/abs/2503.21717