M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ansari, Abolfazl, Zhang, Delvin Ce, Zou, Zhuoyang, Yin, Wenpeng, Lee, Dongwon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915942827556864
author Ansari, Abolfazl
Zhang, Delvin Ce
Zou, Zhuoyang
Yin, Wenpeng
Lee, Dongwon
author_facet Ansari, Abolfazl
Zhang, Delvin Ce
Zou, Zhuoyang
Yin, Wenpeng
Lee, Dongwon
contents Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing benchmarks lack the scale, domain diversity, and visual complexity needed to evaluate this alignment realistically. To address this gap, we introduce M2-Verify, a large-scale multimodal dataset for checking scientific claim consistency. Sourced from PubMed and arXiv, M2-Verify provides over 469K instances across 16 domains, rigorously validated through expert audits. Extensive baseline experiments show that state-of-the-art models struggle to maintain robust consistency. While top models achieve up to 85.8\% Micro-F1 on low-complexity medical perturbations, performance drops to 61.6\% on high-complexity challenges like anatomical shifts. Furthermore, expert evaluations expose hallucinations when models generate scientific explanations for their alignment decisions. Finally, we demonstrate our dataset's utility and provide comprehensive usage guidelines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01306
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
Ansari, Abolfazl
Zhang, Delvin Ce
Zou, Zhuoyang
Yin, Wenpeng
Lee, Dongwon
Computation and Language
Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing benchmarks lack the scale, domain diversity, and visual complexity needed to evaluate this alignment realistically. To address this gap, we introduce M2-Verify, a large-scale multimodal dataset for checking scientific claim consistency. Sourced from PubMed and arXiv, M2-Verify provides over 469K instances across 16 domains, rigorously validated through expert audits. Extensive baseline experiments show that state-of-the-art models struggle to maintain robust consistency. While top models achieve up to 85.8\% Micro-F1 on low-complexity medical perturbations, performance drops to 61.6\% on high-complexity challenges like anatomical shifts. Furthermore, expert evaluations expose hallucinations when models generate scientific explanations for their alignment decisions. Finally, we demonstrate our dataset's utility and provide comprehensive usage guidelines.
title M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
topic Computation and Language
url https://arxiv.org/abs/2604.01306