Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tong, Bingkui, Xia, Jiaer, Zhou, Kaiyang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911183928295424
author Tong, Bingkui
Xia, Jiaer
Zhou, Kaiyang
author_facet Tong, Bingkui
Xia, Jiaer
Zhou, Kaiyang
contents Multimodal Large Language Models (MLLMs) have shown impressive perception and reasoning capabilities, yet they often suffer from hallucinations -- generating outputs that are linguistically coherent but inconsistent with the context of the input image, including inaccuracies in objects, attributes, and relations. To address this challenge, we propose a simple approach called Layer Contrastive Decoding (LayerCD). Our design is motivated by the observation that shallow visual features are much more likely than deep visual features to cause an MLLM to hallucinate as they only capture biased, low-level information that is insufficient for high-level reasoning. Therefore, LayerCD aims to filter out hallucinations by contrasting the output distributions generated from visual features of different levels, specifically those from the shallow and deep layers of the vision encoder, respectively. We conduct extensive experiments on two hallucination benchmarks and show that LayerCD significantly outperforms current state-of-the-art. The code for LayerCD is available at https://github.com/maifoundations/LayerCD .
format Preprint
id arxiv_https___arxiv_org_abs_2509_25177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
Tong, Bingkui
Xia, Jiaer
Zhou, Kaiyang
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have shown impressive perception and reasoning capabilities, yet they often suffer from hallucinations -- generating outputs that are linguistically coherent but inconsistent with the context of the input image, including inaccuracies in objects, attributes, and relations. To address this challenge, we propose a simple approach called Layer Contrastive Decoding (LayerCD). Our design is motivated by the observation that shallow visual features are much more likely than deep visual features to cause an MLLM to hallucinate as they only capture biased, low-level information that is insufficient for high-level reasoning. Therefore, LayerCD aims to filter out hallucinations by contrasting the output distributions generated from visual features of different levels, specifically those from the shallow and deep layers of the vision encoder, respectively. We conduct extensive experiments on two hallucination benchmarks and show that LayerCD significantly outperforms current state-of-the-art. The code for LayerCD is available at https://github.com/maifoundations/LayerCD .
title Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25177