Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Guangtao, Cheng, Xinyi, Xu, Chenghao, Liu, Qi, Yang, Muli, Fang, Fen, Chen, Huilin, Yan, Jiexi, Yang, Xu, Deng, Cheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915689926754304
author Lyu, Guangtao
Cheng, Xinyi
Xu, Chenghao
Liu, Qi
Yang, Muli
Fang, Fen
Chen, Huilin
Yan, Jiexi
Yang, Xu
Deng, Cheng
author_facet Lyu, Guangtao
Cheng, Xinyi
Xu, Chenghao
Liu, Qi
Yang, Muli
Fang, Fen
Chen, Huilin
Yan, Jiexi
Yang, Xu
Deng, Cheng
contents Large Vision-Language Models (LVLMs) have shown remarkable capabilities, yet hallucinations remain a persistent challenge. This work presents a systematic analysis of the internal evolution of visual perception and token generation in LVLMs, revealing two key patterns. First, perception follows a three-stage GATE process: early layers perform a Global scan, intermediate layers Approach and Tighten on core content, and later layers Explore supplementary regions. Second, generation exhibits an SAD (Subdominant Accumulation to Dominant) pattern, where hallucinated tokens arise from the repeated accumulation of subdominant tokens lacking support from attention (visual perception) or feed-forward network (internal knowledge). Guided by these findings, we devise the VDC (Validated Dominance Correction) strategy, which detects unsupported tokens and replaces them with validated dominant ones to improve output reliability. Extensive experiments across multiple models and benchmarks confirm that VDC substantially mitigates hallucinations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
Lyu, Guangtao
Cheng, Xinyi
Xu, Chenghao
Liu, Qi
Yang, Muli
Fang, Fen
Chen, Huilin
Yan, Jiexi
Yang, Xu
Deng, Cheng
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have shown remarkable capabilities, yet hallucinations remain a persistent challenge. This work presents a systematic analysis of the internal evolution of visual perception and token generation in LVLMs, revealing two key patterns. First, perception follows a three-stage GATE process: early layers perform a Global scan, intermediate layers Approach and Tighten on core content, and later layers Explore supplementary regions. Second, generation exhibits an SAD (Subdominant Accumulation to Dominant) pattern, where hallucinated tokens arise from the repeated accumulation of subdominant tokens lacking support from attention (visual perception) or feed-forward network (internal knowledge). Guided by these findings, we devise the VDC (Validated Dominance Correction) strategy, which detects unsupported tokens and replaces them with validated dominant ones to improve output reliability. Extensive experiments across multiple models and benchmarks confirm that VDC substantially mitigates hallucinations.
title Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18813