Multimodal Claim Extraction for Fact-Checking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Teo, Joycelyn, Cao, Rui, Deng, Zhenyun, Ding, Zifeng, Schlichtkrull, Michael Sejr, Vlachos, Andreas
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915942930317312
author Teo, Joycelyn
Cao, Rui
Deng, Zhenyun
Ding, Zifeng
Schlichtkrull, Michael Sejr
Vlachos, Andreas
author_facet Teo, Joycelyn
Cao, Rui
Deng, Zhenyun
Ding, Zifeng
Schlichtkrull, Michael Sejr
Vlachos, Andreas
contents Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with images such as memes, screenshots, and photos, creating challenges that differ from both text-only claim extraction and well-studied multimodal tasks like image captioning or visual question answering. In this work, we present the first benchmark for multimodal claim extraction from social media, consisting of posts containing text and one or more images, annotated with gold-standard claims derived from real-world fact-checkers. We evaluate state-of-the-art multimodal LLMs (MLLMs) under a three-part evaluation framework (semantic alignment, faithfulness, and decontextualization) and find that baseline MLLMs struggle to model rhetorical intent and contextual cues. To address this, we introduce MICE, an intent-aware framework which shows improvements in intent-critical cases.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16311
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Claim Extraction for Fact-Checking
Teo, Joycelyn
Cao, Rui
Deng, Zhenyun
Ding, Zifeng
Schlichtkrull, Michael Sejr
Vlachos, Andreas
Computation and Language
Artificial Intelligence
Social and Information Networks
Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with images such as memes, screenshots, and photos, creating challenges that differ from both text-only claim extraction and well-studied multimodal tasks like image captioning or visual question answering. In this work, we present the first benchmark for multimodal claim extraction from social media, consisting of posts containing text and one or more images, annotated with gold-standard claims derived from real-world fact-checkers. We evaluate state-of-the-art multimodal LLMs (MLLMs) under a three-part evaluation framework (semantic alignment, faithfulness, and decontextualization) and find that baseline MLLMs struggle to model rhetorical intent and contextual cues. To address this, we introduce MICE, an intent-aware framework which shows improvements in intent-critical cases.
title Multimodal Claim Extraction for Fact-Checking
topic Computation and Language
Artificial Intelligence
Social and Information Networks
url https://arxiv.org/abs/2604.16311