Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen, Tuan Dung, Ho, Minh Khoi, Chen, Qi, Xie, Yutong, Cam-Tu, Nguyen, Nguyen, Minh Khoi, Nguyen, Dang Huy Pham, Hengel, Anton van den, Verjans, Johan W., Nguyen, Phi Le, Phan, Vu Minh Hieu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917386599268352
author Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Xie, Yutong
Cam-Tu, Nguyen
Nguyen, Minh Khoi
Nguyen, Dang Huy Pham
Hengel, Anton van den
Verjans, Johan W.
Nguyen, Phi Le
Phan, Vu Minh Hieu
author_facet Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Xie, Yutong
Cam-Tu, Nguyen
Nguyen, Minh Khoi
Nguyen, Dang Huy Pham
Hengel, Anton van den
Verjans, Johan W.
Nguyen, Phi Le
Phan, Vu Minh Hieu
contents Large vision-language models (LVLMs) achieve strong performance on visual reasoning tasks but remain highly susceptible to hallucination. Existing detection methods predominantly rely on coarse, whole-image measures of how an object token relates to the input image. This global strategy is limited: hallucinated tokens may exhibit weak but widely scattered correlations across many local regions, which aggregate into deceptively high overall relevance, thus evading the current global hallucination detectors. We begin with a simple yet critical observation: a faithful object token must be strongly grounded in a specific image region. Building on this insight, we introduce a patch-level hallucination detection framework that examines fine-grained token-level interactions across model layers. Our analysis uncovers two characteristic signatures of hallucinated tokens: (i) they yield diffuse, non-localized attention patterns, in contrast to the compact, well-focused attention seen in faithful tokens; and (ii) they fail to exhibit meaningful semantic alignment with any visual region. Guided by these findings, we develop a lightweight and interpretable detection method that leverages patch-level statistical features, combined with hidden-layer representations. Our approach achieves up to 90% accuracy in token-level hallucination detection, demonstrating the superiority of fine-grained structural analysis for detecting hallucinations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Xie, Yutong
Cam-Tu, Nguyen
Nguyen, Minh Khoi
Nguyen, Dang Huy Pham
Hengel, Anton van den
Verjans, Johan W.
Nguyen, Phi Le
Phan, Vu Minh Hieu
Computer Vision and Pattern Recognition
Large vision-language models (LVLMs) achieve strong performance on visual reasoning tasks but remain highly susceptible to hallucination. Existing detection methods predominantly rely on coarse, whole-image measures of how an object token relates to the input image. This global strategy is limited: hallucinated tokens may exhibit weak but widely scattered correlations across many local regions, which aggregate into deceptively high overall relevance, thus evading the current global hallucination detectors. We begin with a simple yet critical observation: a faithful object token must be strongly grounded in a specific image region. Building on this insight, we introduce a patch-level hallucination detection framework that examines fine-grained token-level interactions across model layers. Our analysis uncovers two characteristic signatures of hallucinated tokens: (i) they yield diffuse, non-localized attention patterns, in contrast to the compact, well-focused attention seen in faithful tokens; and (ii) they fail to exhibit meaningful semantic alignment with any visual region. Guided by these findings, we develop a lightweight and interpretable detection method that leverages patch-level statistical features, combined with hidden-layer representations. Our approach achieves up to 90% accuracy in token-level hallucination detection, demonstrating the superiority of fine-grained structural analysis for detecting hallucinations.
title Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04863