Towards Explainable Bilingual Multimodal Misinformation Detection and Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yiwei, Huang, Zhenglin, Wen, Haiquan, Li, Tianxiao, Dong, Yi, Fei, Hao, Wu, Baoyuan, Cheng, Guangliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912754707726336
author He, Yiwei
Huang, Zhenglin
Wen, Haiquan
Li, Tianxiao
Dong, Yi
Fei, Hao
Wu, Baoyuan
Cheng, Guangliang
author_facet He, Yiwei
Huang, Zhenglin
Wen, Haiquan
Li, Tianxiao
Dong, Yi
Fei, Hao
Wu, Baoyuan
Cheng, Guangliang
contents The increasing realism of multimodal content has made misinformation more subtle and harder to detect, especially in news media where images are frequently paired with bilingual (e.g., Chinese-English) subtitles. Such content often includes localized image edits and cross-lingual inconsistencies that jointly distort meaning while remaining superficially plausible. We introduce BiMi, a bilingual multimodal framework that jointly performs region-level localization, cross-modal and cross-lingual consistency detection, and natural language explanation for misinformation analysis. To support generalization, BiMi integrates an online retrieval module that supplements model reasoning with up-to-date external context. We further release BiMiBench, a large-scale and comprehensive benchmark constructed by systematically editing real news images and subtitles, comprising 104,000 samples with realistic manipulations across visual and linguistic modalities. To enhance interpretability, we apply Group Relative Policy Optimization (GRPO) to improve explanation quality, marking the first use of GRPO in this domain. Extensive experiments demonstrate that BiMi outperforms strong baselines by up to +8.9 in classification accuracy, +15.9 in localization accuracy, and +2.5 in explanation BERTScore, advancing state-of-the-art performance in realistic, multilingual misinformation detection. Code, models, and datasets will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22930
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Explainable Bilingual Multimodal Misinformation Detection and Localization
He, Yiwei
Huang, Zhenglin
Wen, Haiquan
Li, Tianxiao
Dong, Yi
Fei, Hao
Wu, Baoyuan
Cheng, Guangliang
Computer Vision and Pattern Recognition
The increasing realism of multimodal content has made misinformation more subtle and harder to detect, especially in news media where images are frequently paired with bilingual (e.g., Chinese-English) subtitles. Such content often includes localized image edits and cross-lingual inconsistencies that jointly distort meaning while remaining superficially plausible. We introduce BiMi, a bilingual multimodal framework that jointly performs region-level localization, cross-modal and cross-lingual consistency detection, and natural language explanation for misinformation analysis. To support generalization, BiMi integrates an online retrieval module that supplements model reasoning with up-to-date external context. We further release BiMiBench, a large-scale and comprehensive benchmark constructed by systematically editing real news images and subtitles, comprising 104,000 samples with realistic manipulations across visual and linguistic modalities. To enhance interpretability, we apply Group Relative Policy Optimization (GRPO) to improve explanation quality, marking the first use of GRPO in this domain. Extensive experiments demonstrate that BiMi outperforms strong baselines by up to +8.9 in classification accuracy, +15.9 in localization accuracy, and +2.5 in explanation BERTScore, advancing state-of-the-art performance in realistic, multilingual misinformation detection. Code, models, and datasets will be released.
title Towards Explainable Bilingual Multimodal Misinformation Detection and Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.22930