Mitigating Object Hallucination via Robust Local Perception Search

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Zixian, Yang, Chao, Zhou, Zhanhui, Xu, Xing, Lu, Chaochao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910994614190080
author Gao, Zixian
Yang, Chao
Zhou, Zhanhui
Xu, Xing
Lu, Chaochao
author_facet Gao, Zixian
Yang, Chao
Zhou, Zhanhui
Xu, Xing
Lu, Chaochao
contents Recent advancements in Multimodal Large Language Models (MLLMs) have enabled them to effectively integrate vision and language, addressing a variety of downstream tasks. However, despite their significant success, these models still exhibit hallucination phenomena, where the outputs appear plausible but do not align with the content of the images. To mitigate this issue, we introduce Local Perception Search (LPS), a decoding method during inference that is both simple and training-free, yet effectively suppresses hallucinations. This method leverages local visual prior information as a value function to correct the decoding process. Additionally, we observe that the impact of the local visual prior on model performance is more pronounced in scenarios with high levels of image noise. Notably, LPS is a plug-and-play approach that is compatible with various models. Extensive experiments on widely used hallucination benchmarks and noisy data demonstrate that LPS significantly reduces the incidence of hallucinations compared to the baseline, showing exceptional performance, particularly in noisy settings.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06729
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Object Hallucination via Robust Local Perception Search
Gao, Zixian
Yang, Chao
Zhou, Zhanhui
Xu, Xing
Lu, Chaochao
Computer Vision and Pattern Recognition
Computation and Language
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled them to effectively integrate vision and language, addressing a variety of downstream tasks. However, despite their significant success, these models still exhibit hallucination phenomena, where the outputs appear plausible but do not align with the content of the images. To mitigate this issue, we introduce Local Perception Search (LPS), a decoding method during inference that is both simple and training-free, yet effectively suppresses hallucinations. This method leverages local visual prior information as a value function to correct the decoding process. Additionally, we observe that the impact of the local visual prior on model performance is more pronounced in scenarios with high levels of image noise. Notably, LPS is a plug-and-play approach that is compatible with various models. Extensive experiments on widely used hallucination benchmarks and noisy data demonstrate that LPS significantly reduces the incidence of hallucinations compared to the baseline, showing exceptional performance, particularly in noisy settings.
title Mitigating Object Hallucination via Robust Local Perception Search
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.06729