Multimodal Claim Extraction for Fact-Checking
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915942930317312 |
|---|---|
| author | Teo, Joycelyn Cao, Rui Deng, Zhenyun Ding, Zifeng Schlichtkrull, Michael Sejr Vlachos, Andreas |
| author_facet | Teo, Joycelyn Cao, Rui Deng, Zhenyun Ding, Zifeng Schlichtkrull, Michael Sejr Vlachos, Andreas |
| contents | Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with images such as memes, screenshots, and photos, creating challenges that differ from both text-only claim extraction and well-studied multimodal tasks like image captioning or visual question answering. In this work, we present the first benchmark for multimodal claim extraction from social media, consisting of posts containing text and one or more images, annotated with gold-standard claims derived from real-world fact-checkers. We evaluate state-of-the-art multimodal LLMs (MLLMs) under a three-part evaluation framework (semantic alignment, faithfulness, and decontextualization) and find that baseline MLLMs struggle to model rhetorical intent and contextual cues. To address this, we introduce MICE, an intent-aware framework which shows improvements in intent-critical cases. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_16311 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Multimodal Claim Extraction for Fact-Checking Teo, Joycelyn Cao, Rui Deng, Zhenyun Ding, Zifeng Schlichtkrull, Michael Sejr Vlachos, Andreas Computation and Language Artificial Intelligence Social and Information Networks Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with images such as memes, screenshots, and photos, creating challenges that differ from both text-only claim extraction and well-studied multimodal tasks like image captioning or visual question answering. In this work, we present the first benchmark for multimodal claim extraction from social media, consisting of posts containing text and one or more images, annotated with gold-standard claims derived from real-world fact-checkers. We evaluate state-of-the-art multimodal LLMs (MLLMs) under a three-part evaluation framework (semantic alignment, faithfulness, and decontextualization) and find that baseline MLLMs struggle to model rhetorical intent and contextual cues. To address this, we introduce MICE, an intent-aware framework which shows improvements in intent-critical cases. |
| title | Multimodal Claim Extraction for Fact-Checking |
| topic | Computation and Language Artificial Intelligence Social and Information Networks |
| url | https://arxiv.org/abs/2604.16311 |