A Survey of Multimodal Hallucination Evaluation and Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Zhiyuan, Min, Yuecong, Zhang, Jie, Yan, Bei, Wang, Jiahao, Wang, Xiaozhen, Shan, Shiguang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918280387624960
author Chen, Zhiyuan
Min, Yuecong
Zhang, Jie
Yan, Bei
Wang, Jiahao
Wang, Xiaozhen
Shan, Shiguang
author_facet Chen, Zhiyuan
Min, Yuecong
Zhang, Jie
Yan, Bei
Wang, Jiahao
Wang, Xiaozhen
Shan, Shiguang
contents Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However, these models often suffer from hallucination, producing content that appears plausible but contradicts the input content or established world knowledge. This survey offers an in-depth review of hallucination evaluation benchmarks and detection methods across Image-to-Text (I2T) and Text-to-image (T2I) generation tasks. Specifically, we first propose a taxonomy of hallucination based on faithfulness and factuality, incorporating the common types of hallucinations observed in practice. Then we provide an overview of existing hallucination evaluation benchmarks for both T2I and I2T tasks, highlighting their construction process, evaluation objectives, and employed metrics. Furthermore, we summarize recent advances in hallucination detection methods, which aims to identify hallucinated content at the instance level and serve as a practical complement of benchmark-based evaluation. Finally, we highlight key limitations in current benchmarks and detection methods, and outline potential directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey of Multimodal Hallucination Evaluation and Detection
Chen, Zhiyuan
Min, Yuecong
Zhang, Jie
Yan, Bei
Wang, Jiahao
Wang, Xiaozhen
Shan, Shiguang
Computer Vision and Pattern Recognition
Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However, these models often suffer from hallucination, producing content that appears plausible but contradicts the input content or established world knowledge. This survey offers an in-depth review of hallucination evaluation benchmarks and detection methods across Image-to-Text (I2T) and Text-to-image (T2I) generation tasks. Specifically, we first propose a taxonomy of hallucination based on faithfulness and factuality, incorporating the common types of hallucinations observed in practice. Then we provide an overview of existing hallucination evaluation benchmarks for both T2I and I2T tasks, highlighting their construction process, evaluation objectives, and employed metrics. Furthermore, we summarize recent advances in hallucination detection methods, which aims to identify hallucinated content at the instance level and serve as a practical complement of benchmark-based evaluation. Finally, we highlight key limitations in current benchmarks and detection methods, and outline potential directions for future research.
title A Survey of Multimodal Hallucination Evaluation and Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19024