Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zeng, Fengzhu, Li, Wenqian, Gao, Wei, Pang, Yan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909329133666304
author Zeng, Fengzhu
Li, Wenqian
Gao, Wei
Pang, Yan
author_facet Zeng, Fengzhu
Li, Wenqian
Gao, Wei
Pang, Yan
contents Detecting multimodal misinformation, especially in the form of image-text pairs, is crucial. Obtaining large-scale, high-quality real-world fact-checking datasets for training detectors is costly, leading researchers to use synthetic datasets generated by AI technologies. However, the generalizability of detectors trained on synthetic data to real-world scenarios remains unclear due to the distribution gap. To address this, we propose learning from synthetic data for detecting real-world multimodal misinformation through two model-agnostic data selection methods that match synthetic and real-world data distributions. Experiments show that our method enhances the performance of a small MLLM (13B) on real-world fact-checking datasets, enabling it to even surpass GPT-4V~\cite{GPT-4V}.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19656
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs
Zeng, Fengzhu
Li, Wenqian
Gao, Wei
Pang, Yan
Computation and Language
Detecting multimodal misinformation, especially in the form of image-text pairs, is crucial. Obtaining large-scale, high-quality real-world fact-checking datasets for training detectors is costly, leading researchers to use synthetic datasets generated by AI technologies. However, the generalizability of detectors trained on synthetic data to real-world scenarios remains unclear due to the distribution gap. To address this, we propose learning from synthetic data for detecting real-world multimodal misinformation through two model-agnostic data selection methods that match synthetic and real-world data distributions. Experiments show that our method enhances the performance of a small MLLM (13B) on real-world fact-checking datasets, enabling it to even surpass GPT-4V~\cite{GPT-4V}.
title Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs
topic Computation and Language
url https://arxiv.org/abs/2409.19656