Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Jiahe, He, Jiaying, Chen, Qiyuan, Shao, Qian, Ying, Jiahe, Xu, Hongxia, Chen, Jintai, Zheng, Jianwei, Wu, Jian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910053282349056
author Chen, Jiahe
He, Jiaying
Chen, Qiyuan
Shao, Qian
Ying, Jiahe
Xu, Hongxia
Chen, Jintai
Zheng, Jianwei
Wu, Jian
author_facet Chen, Jiahe
He, Jiaying
Chen, Qiyuan
Shao, Qian
Ying, Jiahe
Xu, Hongxia
Chen, Jintai
Zheng, Jianwei
Wu, Jian
contents Large Vision-Language Models (LVLMs) face a tug-of-war between powerful linguistic priors and visual evidence, often leading to \emph{semantic drift}: a progressive detachment from the input image that can abruptly emerge at specific decoding steps. Through a token-level diagnosis, we show that hallucination is frequently triggered not by the absence of grounded candidates, but by a failure of selection -- the model chooses a linguistically convenient yet visually unfaithful token even when better grounded alternatives exist. Motivated by this insight, we propose \textbf{D}ynamic \textbf{L}ogits \textbf{C}alibration (DLC), a training-free decoding framework that introduces a lightweight visual referee to intervene exactly when drift happens. At each step, DLC performs a dual-aspect grounding check on top-$k$ candidates: (1) it assesses the intrinsic visual relevance of a candidate token and (2) its contextual visual coherence. These signals are evaluated against an adaptive historical baseline to compute a relative visual advantage, which is then used to dynamically calibrate logits and favor grounded tokens. Extensive experiments on CHAIR, POPE, SHR, GPT-4o evaluation, and MME demonstrate that DLC consistently reduces hallucinations across multiple LVLMs while preserving response quality. Further analyses validate robustness to different vision backbones and demonstrate a favorable trade-off between output quality and computational cost as the candidate pool size varies. Code will be released on https://github.com/JiaheChen2002/DLC.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
Chen, Jiahe
He, Jiaying
Chen, Qiyuan
Shao, Qian
Ying, Jiahe
Xu, Hongxia
Chen, Jintai
Zheng, Jianwei
Wu, Jian
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) face a tug-of-war between powerful linguistic priors and visual evidence, often leading to \emph{semantic drift}: a progressive detachment from the input image that can abruptly emerge at specific decoding steps. Through a token-level diagnosis, we show that hallucination is frequently triggered not by the absence of grounded candidates, but by a failure of selection -- the model chooses a linguistically convenient yet visually unfaithful token even when better grounded alternatives exist. Motivated by this insight, we propose \textbf{D}ynamic \textbf{L}ogits \textbf{C}alibration (DLC), a training-free decoding framework that introduces a lightweight visual referee to intervene exactly when drift happens. At each step, DLC performs a dual-aspect grounding check on top-$k$ candidates: (1) it assesses the intrinsic visual relevance of a candidate token and (2) its contextual visual coherence. These signals are evaluated against an adaptive historical baseline to compute a relative visual advantage, which is then used to dynamically calibrate logits and favor grounded tokens. Extensive experiments on CHAIR, POPE, SHR, GPT-4o evaluation, and MME demonstrate that DLC consistently reduces hallucinations across multiple LVLMs while preserving response quality. Further analyses validate robustness to different vision backbones and demonstrate a favorable trade-off between output quality and computational cost as the candidate pool size varies. Code will be released on https://github.com/JiaheChen2002/DLC.
title Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21509