DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Weilin, Ying, Zonghao, Meng, Chunlei, Liu, Jiahui, Zhou, Hengyang, Zou, Quanchen, Zhang, Deyue, Yang, Dongdong, Zhang, Xiangzheng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915722755571712
author Zhou, Weilin
Ying, Zonghao
Meng, Chunlei
Liu, Jiahui
Zhou, Hengyang
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
author_facet Zhou, Weilin
Ying, Zonghao
Meng, Chunlei
Liu, Jiahui
Zhou, Hengyang
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
contents Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundations. To address this, we propose DIVER (Dynamic Iterative Visual Evidence Reasoning), a framework grounded in a progressive, evidence-driven reasoning paradigm. DIVER first establishes a strong text-based baseline through language analysis, leveraging intra-modal consistency to filter unreliable or hallucinated claims. Only when textual evidence is insufficient does the framework introduce visual information, where inter-modal alignment verification adaptively determines whether deeper visual inspection is necessary. For samples exhibiting significant cross-modal semantic discrepancies, DIVER selectively invokes fine-grained visual tools (e.g., OCR and dense captioning) to extract task-relevant evidence, which is iteratively aggregated via uncertainty-aware fusion to refine multimodal reasoning. Experiments on Weibo, Weibo21, and GossipCop demonstrate that DIVER outperforms state-of-the-art baselines by an average of 2.72\%, while optimizing inference efficiency with a reduced latency of 4.12 s.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07178
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
Zhou, Weilin
Ying, Zonghao
Meng, Chunlei
Liu, Jiahui
Zhou, Hengyang
Zou, Quanchen
Zhang, Deyue
Yang, Dongdong
Zhang, Xiangzheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundations. To address this, we propose DIVER (Dynamic Iterative Visual Evidence Reasoning), a framework grounded in a progressive, evidence-driven reasoning paradigm. DIVER first establishes a strong text-based baseline through language analysis, leveraging intra-modal consistency to filter unreliable or hallucinated claims. Only when textual evidence is insufficient does the framework introduce visual information, where inter-modal alignment verification adaptively determines whether deeper visual inspection is necessary. For samples exhibiting significant cross-modal semantic discrepancies, DIVER selectively invokes fine-grained visual tools (e.g., OCR and dense captioning) to extract task-relevant evidence, which is iteratively aggregated via uncertainty-aware fusion to refine multimodal reasoning. Experiments on Weibo, Weibo21, and GossipCop demonstrate that DIVER outperforms state-of-the-art baselines by an average of 2.72\%, while optimizing inference efficiency with a reduced latency of 4.12 s.
title DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.07178