Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Bing, Li, Ximing, Wang, Yanjun, Li, Changchun, Wu, Lin Yuanbo, Wang, Buyu, Wang, Shengsheng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908639122423808
author Wang, Bing
Li, Ximing
Wang, Yanjun
Li, Changchun
Wu, Lin Yuanbo
Wang, Buyu
Wang, Shengsheng
author_facet Wang, Bing
Li, Ximing
Wang, Yanjun
Li, Changchun
Wu, Lin Yuanbo
Wang, Buyu
Wang, Shengsheng
contents Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality because the text generally describes the whole event/story of the current post but the image often presents partial scenes only. Our preliminary empirical results indicate that the image modality exactly contributes less to MMD. Upon this idea, we propose a new MMD method named RETSIMD. Specifically, we suppose that each text can be divided into several segments, and each text segment describes a partial scene that can be presented by an image. Accordingly, we split the text into a sequence of segments, and feed these segments into a pre-trained text-to-image generator to augment a sequence of images. We further incorporate two auxiliary objectives concerning text-image and image-label mutual information, and further post-train the generator over an auxiliary text-to-image generation benchmark dataset. Additionally, we propose a graph structure by defining three heuristic relationships between images, and use a graph neural network to generate the fused features. Extensive empirical results validate the effectiveness of RETSIMD.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
Wang, Bing
Li, Ximing
Wang, Yanjun
Li, Changchun
Wu, Lin Yuanbo
Wang, Buyu
Wang, Shengsheng
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality because the text generally describes the whole event/story of the current post but the image often presents partial scenes only. Our preliminary empirical results indicate that the image modality exactly contributes less to MMD. Upon this idea, we propose a new MMD method named RETSIMD. Specifically, we suppose that each text can be divided into several segments, and each text segment describes a partial scene that can be presented by an image. Accordingly, we split the text into a sequence of segments, and feed these segments into a pre-trained text-to-image generator to augment a sequence of images. We further incorporate two auxiliary objectives concerning text-image and image-label mutual information, and further post-train the generator over an auxiliary text-to-image generation benchmark dataset. Additionally, we propose a graph structure by defining three heuristic relationships between images, and use a graph neural network to generate the fused features. Extensive empirical results validate the effectiveness of RETSIMD.
title Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2511.06284