SciClaimEval: Cross-modal Claim Verification in Scientific Papers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ho, Xanh, Wu, Yun-Ang, Kumar, Sunisth, Xia, Tian Cheng, Boudin, Florian, Greiner-Petter, Andre, Aizawa, Akiko
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918336483295232
author Ho, Xanh
Wu, Yun-Ang
Kumar, Sunisth
Xia, Tian Cheng
Boudin, Florian
Greiner-Petter, Andre
Aizawa, Akiko
author_facet Ho, Xanh
Wu, Yun-Ang
Kumar, Sunisth
Xia, Tian Cheng
Boudin, Florian
Greiner-Petter, Andre
Aizawa, Akiko
contents We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, directly extracted from published papers. To create refuted claims, we introduce a novel approach that modifies the supporting evidence (figures and tables), rather than altering the claims or relying on large language models (LLMs) to fabricate contradictions. The dataset provides cross-modal evidence with diverse representations: figures are available as images, while tables are provided in multiple formats, including images, LaTeX source, HTML, and JSON. SciClaimEval contains 1,664 annotated samples from 180 papers across three domains, machine learning, natural language processing, and medicine, validated through expert annotation. We benchmark 11 multimodal foundation models, both open-source and proprietary, across the dataset. Results show that figure-based verification remains particularly challenging for all models, as a substantial performance gap remains between the best system and human baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07621
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Ho, Xanh
Wu, Yun-Ang
Kumar, Sunisth
Xia, Tian Cheng
Boudin, Florian
Greiner-Petter, Andre
Aizawa, Akiko
Computation and Language
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, directly extracted from published papers. To create refuted claims, we introduce a novel approach that modifies the supporting evidence (figures and tables), rather than altering the claims or relying on large language models (LLMs) to fabricate contradictions. The dataset provides cross-modal evidence with diverse representations: figures are available as images, while tables are provided in multiple formats, including images, LaTeX source, HTML, and JSON. SciClaimEval contains 1,664 annotated samples from 180 papers across three domains, machine learning, natural language processing, and medicine, validated through expert annotation. We benchmark 11 multimodal foundation models, both open-source and proprietary, across the dataset. Results show that figure-based verification remains particularly challenging for all models, as a substantial performance gap remains between the best system and human baseline.
title SciClaimEval: Cross-modal Claim Verification in Scientific Papers
topic Computation and Language
url https://arxiv.org/abs/2602.07621