Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Ye, Wang, Yunan, Yu, Zitong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915282709118976
author Zhu, Ye
Wang, Yunan
Yu, Zitong
author_facet Zhu, Ye
Wang, Yunan
Yu, Zitong
contents Multimodal news contains a wealth of information and is easily affected by deepfake modeling attacks. To combat the latest image and text generation methods, we present a new Multimodal Fake News Detection dataset (MFND) containing 11 manipulated types, designed to detect and localize highly authentic fake news. Furthermore, we propose a Shallow-Deep Multitask Learning (SDML) model for fake news, which fully uses unimodal and mutual modal features to mine the intrinsic semantics of news. Under shallow inference, we propose the momentum distillation-based light punishment contrastive learning for fine-grained uniform spatial image and text semantic alignment, and an adaptive cross-modal fusion module to enhance mutual modal features. Under deep inference, we design a two-branch framework to augment the image and text unimodal features, respectively merging with mutual modalities features, for four predictions via dedicated detection and localization projections. Experiments on both mainstream and our proposed datasets demonstrate the superiority of the model. Codes and dataset are released at https://github.com/yunan-wang33/sdml.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
Zhu, Ye
Wang, Yunan
Yu, Zitong
Computer Vision and Pattern Recognition
Multimodal news contains a wealth of information and is easily affected by deepfake modeling attacks. To combat the latest image and text generation methods, we present a new Multimodal Fake News Detection dataset (MFND) containing 11 manipulated types, designed to detect and localize highly authentic fake news. Furthermore, we propose a Shallow-Deep Multitask Learning (SDML) model for fake news, which fully uses unimodal and mutual modal features to mine the intrinsic semantics of news. Under shallow inference, we propose the momentum distillation-based light punishment contrastive learning for fine-grained uniform spatial image and text semantic alignment, and an adaptive cross-modal fusion module to enhance mutual modal features. Under deep inference, we design a two-branch framework to augment the image and text unimodal features, respectively merging with mutual modalities features, for four predictions via dedicated detection and localization projections. Experiments on both mainstream and our proposed datasets demonstrate the superiority of the model. Codes and dataset are released at https://github.com/yunan-wang33/sdml.
title Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.06796