Piecing It All Together: Verifying Multi-Hop Multimodal Claims

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haoran, Rangapur, Aman, Xu, Xiongxiao, Liang, Yueqing, Gharwi, Haroon, Yang, Carl, Shu, Kai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912153952321536
author Wang, Haoran
Rangapur, Aman
Xu, Xiongxiao
Liang, Yueqing
Gharwi, Haroon
Yang, Carl
Shu, Kai
author_facet Wang, Haoran
Rangapur, Aman
Xu, Xiongxiao
Liang, Yueqing
Gharwi, Haroon
Yang, Carl
Shu, Kai
contents Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. To address this, we introduce a new task: multi-hop multimodal claim verification. This task challenges models to reason over multiple pieces of evidence from diverse sources, including text, images, and tables, and determine whether the combined multimodal evidence supports or refutes a given claim. To study this task, we construct MMCV, a large-scale dataset comprising 15k multi-hop claims paired with multimodal evidence, generated and refined using large language models, with additional input from human feedback. We show that MMCV is challenging even for the latest state-of-the-art multimodal large language models, especially as the number of reasoning hops increases. Additionally, we establish a human performance benchmark on a subset of MMCV. We hope this dataset and its evaluation task will encourage future research in multimodal multi-hop claim verification.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Piecing It All Together: Verifying Multi-Hop Multimodal Claims
Wang, Haoran
Rangapur, Aman
Xu, Xiongxiao
Liang, Yueqing
Gharwi, Haroon
Yang, Carl
Shu, Kai
Computation and Language
Artificial Intelligence
Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. To address this, we introduce a new task: multi-hop multimodal claim verification. This task challenges models to reason over multiple pieces of evidence from diverse sources, including text, images, and tables, and determine whether the combined multimodal evidence supports or refutes a given claim. To study this task, we construct MMCV, a large-scale dataset comprising 15k multi-hop claims paired with multimodal evidence, generated and refined using large language models, with additional input from human feedback. We show that MMCV is challenging even for the latest state-of-the-art multimodal large language models, especially as the number of reasoning hops increases. Additionally, we establish a human performance benchmark on a subset of MMCV. We hope this dataset and its evaluation task will encourage future research in multimodal multi-hop claim verification.
title Piecing It All Together: Verifying Multi-Hop Multimodal Claims
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.09547