MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Chenxi, Chen, Xiang, Zhang, Ningyu, Tian, Bozhong, Xu, Haoming, Deng, Shumin, Chen, Huajun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909505014464512
author Wang, Chenxi
Chen, Xiang
Zhang, Ningyu
Tian, Bozhong
Xu, Haoming
Deng, Shumin
Chen, Huajun
author_facet Wang, Chenxi
Chen, Xiang
Zhang, Ningyu
Tian, Bozhong
Xu, Haoming
Deng, Shumin
Chen, Huajun
contents Multimodal Large Language Models (MLLMs) frequently exhibit hallucination phenomena, but the underlying reasons remain poorly understood. In this paper, we present an empirical analysis and find that, although MLLMs incorrectly generate the objects in the final output, they are actually able to recognize visual objects in the preceding layers. We speculate that this may be due to the strong knowledge priors of the language model suppressing the visual information, leading to hallucinations. Motivated by this, we propose a novel dynamic correction decoding method for MLLMs DeCo, which adaptively selects the appropriate preceding layers and proportionally integrates knowledge into the final layer to adjust the output logits. Note that DeCo is model agnostic and can be seamlessly incorporated with various classic decoding strategies and applied to different MLLMs. We evaluate DeCo on widely-used benchmarks, demonstrating that it can reduce hallucination rates by a large margin compared to baselines, highlighting its potential to mitigate hallucinations. Code is available at https://github.com/zjunlp/DeCo.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11779
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
Wang, Chenxi
Chen, Xiang
Zhang, Ningyu
Tian, Bozhong
Xu, Haoming
Deng, Shumin
Chen, Huajun
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Multimodal Large Language Models (MLLMs) frequently exhibit hallucination phenomena, but the underlying reasons remain poorly understood. In this paper, we present an empirical analysis and find that, although MLLMs incorrectly generate the objects in the final output, they are actually able to recognize visual objects in the preceding layers. We speculate that this may be due to the strong knowledge priors of the language model suppressing the visual information, leading to hallucinations. Motivated by this, we propose a novel dynamic correction decoding method for MLLMs DeCo, which adaptively selects the appropriate preceding layers and proportionally integrates knowledge into the final layer to adjust the output logits. Note that DeCo is model agnostic and can be seamlessly incorporated with various classic decoding strategies and applied to different MLLMs. We evaluate DeCo on widely-used benchmarks, demonstrating that it can reduce hallucination rates by a large margin compared to baselines, highlighting its potential to mitigate hallucinations. Code is available at https://github.com/zjunlp/DeCo.
title MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2410.11779