MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Xuannan, Li, Zekun, Li, Peipei, Huang, Huaibo, Xia, Shuhan, Cui, Xing, Huang, Linzhi, Deng, Weihong, He, Zhaofeng
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913708893011968
author Liu, Xuannan
Li, Zekun
Li, Peipei
Huang, Huaibo
Xia, Shuhan
Cui, Xing
Huang, Linzhi
Deng, Weihong
He, Zhaofeng
author_facet Liu, Xuannan
Li, Zekun
Li, Peipei
Huang, Huaibo
Xia, Shuhan
Cui, Xing
Huang, Linzhi
Deng, Weihong
He, Zhaofeng
contents Current multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08772
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
Liu, Xuannan
Li, Zekun
Li, Peipei
Huang, Huaibo
Xia, Shuhan
Cui, Xing
Huang, Linzhi
Deng, Weihong
He, Zhaofeng
Computer Vision and Pattern Recognition
Computation and Language
Current multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods.
title MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2406.08772