Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Meng, Wu, Minghao, Rajan, Deepu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910241134739456
author Shen, Meng
Wu, Minghao
Rajan, Deepu
author_facet Shen, Meng
Wu, Minghao
Rajan, Deepu
contents Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that one possible origin of hallucination is the model's tendency to prioritize text generation over meaningful interaction with images. To explore this, we examine the generation process and categorize text tokens into three groups: image-positive, invariant, and negative, based on their visual dependence on input image tokens. Our analysis reveals that most generated tokens are minimally influenced by the image information. This suggests that during the model's training stage, more emphasis is placed on learning how to follow textual instructions, rather than extracting information from images. Based on this finding, we propose adjusting the training weights of different tokens depending on their visual dependence to control hallucination. Additionally, we remove a portion of the training data that potentially contains more hallucinations as a data filtering strategy. Both methods achieve a reduction in hallucination without compromising response length or introducing additional computational costs during inference. We validate our methods across three LVLM variants, demonstrating the effectiveness and general applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21300
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
Shen, Meng
Wu, Minghao
Rajan, Deepu
Computer Vision and Pattern Recognition
I.2.10; I.4.8
Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that one possible origin of hallucination is the model's tendency to prioritize text generation over meaningful interaction with images. To explore this, we examine the generation process and categorize text tokens into three groups: image-positive, invariant, and negative, based on their visual dependence on input image tokens. Our analysis reveals that most generated tokens are minimally influenced by the image information. This suggests that during the model's training stage, more emphasis is placed on learning how to follow textual instructions, rather than extracting information from images. Based on this finding, we propose adjusting the training weights of different tokens depending on their visual dependence to control hallucination. Additionally, we remove a portion of the training data that potentially contains more hallucinations as a data filtering strategy. Both methods achieve a reduction in hallucination without compromising response length or introducing additional computational costs during inference. We validate our methods across three LVLM variants, demonstrating the effectiveness and general applicability.
title Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
topic Computer Vision and Pattern Recognition
I.2.10; I.4.8
url https://arxiv.org/abs/2605.21300