Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Suo, Wei, Zhang, Lijun, Sun, Mengyang, Wu, Lin Yuanbo, Wang, Peng, Zhang, Yanning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915177728835584
author Suo, Wei
Zhang, Lijun
Sun, Mengyang
Wu, Lin Yuanbo
Wang, Peng
Zhang, Yanning
author_facet Suo, Wei
Zhang, Lijun
Sun, Mengyang
Wu, Lin Yuanbo
Wang, Peng
Zhang, Yanning
contents Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) strategies have been proposed to alleviate hallucination by introducing disturbed inputs. Although great progress has been made, these CD strategies mostly apply a one-size-fits-all approach for all input conditions. In this paper, we revisit this process through extensive experiments. Related results show that hallucination causes are hybrid and each generative step faces a unique hallucination challenge. Leveraging these meaningful insights, we introduce a simple yet effective Octopus-like framework that enables the model to adaptively identify hallucination types and create a dynamic CD workflow. Our Octopus framework not only outperforms existing methods across four benchmarks but also demonstrates excellent deployability and expansibility. Code is available at https://github.com/LijunZhang01/Octopus.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
Suo, Wei
Zhang, Lijun
Sun, Mengyang
Wu, Lin Yuanbo
Wang, Peng
Zhang, Yanning
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) strategies have been proposed to alleviate hallucination by introducing disturbed inputs. Although great progress has been made, these CD strategies mostly apply a one-size-fits-all approach for all input conditions. In this paper, we revisit this process through extensive experiments. Related results show that hallucination causes are hybrid and each generative step faces a unique hallucination challenge. Leveraging these meaningful insights, we introduce a simple yet effective Octopus-like framework that enables the model to adaptively identify hallucination types and create a dynamic CD workflow. Our Octopus framework not only outperforms existing methods across four benchmarks but also demonstrates excellent deployability and expansibility. Code is available at https://github.com/LijunZhang01/Octopus.
title Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.00361