Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Xinyu, Chen, Beitao, Gao, Lianli, Song, Jingkuan, Shen, Heng Tao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913768880996352
author Lyu, Xinyu
Chen, Beitao
Gao, Lianli
Song, Jingkuan
Shen, Heng Tao
author_facet Lyu, Xinyu
Chen, Beitao
Gao, Lianli
Song, Jingkuan
Shen, Heng Tao
contents Although Large Visual Language Models (LVLMs) have demonstrated exceptional abilities in understanding multimodal data, they invariably suffer from hallucinations, leading to a disconnect between the generated text and the corresponding images. Almost all current visual contrastive decoding methods attempt to mitigate these hallucinations by introducing visual uncertainty information that appropriately widens the contrastive logits gap between hallucinatory and targeted ones. However, due to uncontrollable nature of the global visual uncertainty, they struggle to precisely induce the hallucinatory tokens, which severely limits their effectiveness in mitigating hallucinations and may even lead to the generation of undesired hallucinations. To tackle this issue, we conducted the theoretical analysis to promote the effectiveness of contrast decoding. Building on this insight, we introduce a novel optimization strategy named Hallucination-Induced Optimization (HIO). This strategy seeks to amplify the contrast between hallucinatory and targeted tokens relying on a fine-tuned theoretical preference model (i.e., Contrary Bradley-Terry Model), thereby facilitating efficient contrast decoding to alleviate hallucinations in LVLMs. Extensive experimental research demonstrates that our HIO strategy can effectively reduce hallucinations in LVLMs, outperforming state-of-the-art methods across various benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
Lyu, Xinyu
Chen, Beitao
Gao, Lianli
Song, Jingkuan
Shen, Heng Tao
Computer Vision and Pattern Recognition
Although Large Visual Language Models (LVLMs) have demonstrated exceptional abilities in understanding multimodal data, they invariably suffer from hallucinations, leading to a disconnect between the generated text and the corresponding images. Almost all current visual contrastive decoding methods attempt to mitigate these hallucinations by introducing visual uncertainty information that appropriately widens the contrastive logits gap between hallucinatory and targeted ones. However, due to uncontrollable nature of the global visual uncertainty, they struggle to precisely induce the hallucinatory tokens, which severely limits their effectiveness in mitigating hallucinations and may even lead to the generation of undesired hallucinations. To tackle this issue, we conducted the theoretical analysis to promote the effectiveness of contrast decoding. Building on this insight, we introduce a novel optimization strategy named Hallucination-Induced Optimization (HIO). This strategy seeks to amplify the contrast between hallucinatory and targeted tokens relying on a fine-tuned theoretical preference model (i.e., Contrary Bradley-Terry Model), thereby facilitating efficient contrast decoding to alleviate hallucinations in LVLMs. Extensive experimental research demonstrates that our HIO strategy can effectively reduce hallucinations in LVLMs, outperforming state-of-the-art methods across various benchmarks.
title Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15356