Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Duwal, Sharad, Shopnil, Mir Nafis Sharear, Tyagi, Abhishek, Proma, Adiba Mahbub
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913857425899520
author Duwal, Sharad
Shopnil, Mir Nafis Sharear
Tyagi, Abhishek
Proma, Adiba Mahbub
author_facet Duwal, Sharad
Shopnil, Mir Nafis Sharear
Tyagi, Abhishek
Proma, Adiba Mahbub
contents Multimodal out-of-context (OOC) misinformation is misinformation that repurposes real images with unrelated or misleading captions. Detecting such misinformation is challenging because it requires resolving the context of the claim before checking for misinformation. Many current methods, including LLMs and LVLMs, do not perform this contextualization step. LLMs hallucinate in absence of context or parametric knowledge. In this work, we propose a graph-based method that evaluates the consistency between the image and the caption by constructing two graph representations: an evidence graph, derived from online textual evidence, and a claim graph, from the claim in the caption. Using graph neural networks (GNNs) to encode and compare these representations, our framework then evaluates the truthfulness of image-caption pairs. We create datasets for our graph-based method, evaluate and compare our baseline model against popular LLMs on the misinformation detection task. Our method scores $93.05\%$ detection accuracy on the evaluation set and outperforms the second-best performing method (an LLM) by $2.82\%$, making a case for smaller and task-specific methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18221
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs
Duwal, Sharad
Shopnil, Mir Nafis Sharear
Tyagi, Abhishek
Proma, Adiba Mahbub
Machine Learning
Artificial Intelligence
Computation and Language
Information Retrieval
Multimodal out-of-context (OOC) misinformation is misinformation that repurposes real images with unrelated or misleading captions. Detecting such misinformation is challenging because it requires resolving the context of the claim before checking for misinformation. Many current methods, including LLMs and LVLMs, do not perform this contextualization step. LLMs hallucinate in absence of context or parametric knowledge. In this work, we propose a graph-based method that evaluates the consistency between the image and the caption by constructing two graph representations: an evidence graph, derived from online textual evidence, and a claim graph, from the claim in the caption. Using graph neural networks (GNNs) to encode and compare these representations, our framework then evaluates the truthfulness of image-caption pairs. We create datasets for our graph-based method, evaluate and compare our baseline model against popular LLMs on the misinformation detection task. Our method scores $93.05\%$ detection accuracy on the evaluation set and outperforms the second-best performing method (an LLM) by $2.82\%$, making a case for smaller and task-specific methods.
title Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs
topic Machine Learning
Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2505.18221