MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Delvin Ce, Cui, Suhan, Chu, Zhelin, Zhang, Xianren, Lee, Dongwon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912895130927104
author Zhang, Delvin Ce
Cui, Suhan
Chu, Zhelin
Zhang, Xianren
Lee, Dongwon
author_facet Zhang, Delvin Ce
Cui, Suhan
Chu, Zhelin
Zhang, Xianren
Lee, Dongwon
contents Verifying the truthfulness of claims usually requires joint multi-modal reasoning over both textual and visual evidence, such as analyzing both textual caption and chart image for claim verification. In addition, to make the reasoning process transparent, a textual explanation is necessary to justify the verification result. However, most claim verification works mainly focus on the reasoning over textual evidence only or ignore the explainability, resulting in inaccurate and unconvincing verification. To address this problem, we propose a novel model that jointly achieves evidence retrieval, multi-modal claim verification, and explanation generation. For evidence retrieval, we construct a two-layer multi-modal graph for claims and evidence, where we design image-to-text and text-to-image reasoning for multi-modal retrieval. For claim verification, we propose token- and evidence-level fusion to integrate claim and evidence embeddings for multi-modal verification. For explanation generation, we introduce multi-modal Fusion-in-Decoder for explainability. Finally, since almost all the datasets are in general domain, we create a scientific dataset, AIChartClaim, in AI domain to complement claim verification community. Experiments show the strength of our model.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval
Zhang, Delvin Ce
Cui, Suhan
Chu, Zhelin
Zhang, Xianren
Lee, Dongwon
Computation and Language
Verifying the truthfulness of claims usually requires joint multi-modal reasoning over both textual and visual evidence, such as analyzing both textual caption and chart image for claim verification. In addition, to make the reasoning process transparent, a textual explanation is necessary to justify the verification result. However, most claim verification works mainly focus on the reasoning over textual evidence only or ignore the explainability, resulting in inaccurate and unconvincing verification. To address this problem, we propose a novel model that jointly achieves evidence retrieval, multi-modal claim verification, and explanation generation. For evidence retrieval, we construct a two-layer multi-modal graph for claims and evidence, where we design image-to-text and text-to-image reasoning for multi-modal retrieval. For claim verification, we propose token- and evidence-level fusion to integrate claim and evidence embeddings for multi-modal verification. For explanation generation, we introduce multi-modal Fusion-in-Decoder for explainability. Finally, since almost all the datasets are in general domain, we create a scientific dataset, AIChartClaim, in AI domain to complement claim verification community. Experiments show the strength of our model.
title MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval
topic Computation and Language
url https://arxiv.org/abs/2602.10023