Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shoby, Abin, Huy, Ta Duc, Nguyen, Tuan Dung, Ho, Minh Khoi, Chen, Qi, Hengel, Anton van den, Nguyen, Phi Le, Verjans, Johan W., Phan, Vu Minh Hieu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908918737797120
author Shoby, Abin
Huy, Ta Duc
Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Hengel, Anton van den
Nguyen, Phi Le
Verjans, Johan W.
Phan, Vu Minh Hieu
author_facet Shoby, Abin
Huy, Ta Duc
Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Hengel, Anton van den
Nguyen, Phi Le
Verjans, Johan W.
Phan, Vu Minh Hieu
contents Vision Language models (VLMs) often hallucinate non-existent objects. Detecting hallucination is analogous to detecting deception: a single final statement is insufficient, one must examine the underlying reasoning process. Yet existing detectors rely mostly on final-layer signals. Attention-based methods assume hallucinated tokens exhibit low attention, while entropy-based ones use final-step uncertainty. Our analysis reveals the opposite: hallucinated objects can exhibit peaked attention due to contextual priors; and models often express high confidence because intermediate layers have already converged to an incorrect hypothesis. We show that the key to hallucination detection lies within the model's thought process, not its final output. By probing decoder layers, we uncover a previously overlooked behavior, overthinking: models repeatedly revise object hypotheses across layers before committing to an incorrect answer. Once the model latches onto a confounded hypothesis, it can propagate through subsequent layers, ultimately causing hallucination. To capture this behavior, we introduce the Overthinking Score, a metric to measure how many competing hypotheses the model entertains and how unstable these hypotheses are across layers. This score significantly improves hallucination detection: 78.9% F1 on MSCOCO and 71.58% on AMBER.
format Preprint
id arxiv_https___arxiv_org_abs_2603_07619
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
Shoby, Abin
Huy, Ta Duc
Nguyen, Tuan Dung
Ho, Minh Khoi
Chen, Qi
Hengel, Anton van den
Nguyen, Phi Le
Verjans, Johan W.
Phan, Vu Minh Hieu
Computer Vision and Pattern Recognition
Vision Language models (VLMs) often hallucinate non-existent objects. Detecting hallucination is analogous to detecting deception: a single final statement is insufficient, one must examine the underlying reasoning process. Yet existing detectors rely mostly on final-layer signals. Attention-based methods assume hallucinated tokens exhibit low attention, while entropy-based ones use final-step uncertainty. Our analysis reveals the opposite: hallucinated objects can exhibit peaked attention due to contextual priors; and models often express high confidence because intermediate layers have already converged to an incorrect hypothesis. We show that the key to hallucination detection lies within the model's thought process, not its final output. By probing decoder layers, we uncover a previously overlooked behavior, overthinking: models repeatedly revise object hypotheses across layers before committing to an incorrect answer. Once the model latches onto a confounded hypothesis, it can propagate through subsequent layers, ultimately causing hallucination. To capture this behavior, we introduce the Overthinking Score, a metric to measure how many competing hypotheses the model entertains and how unstable these hypotheses are across layers. This score significantly improves hallucination detection: 78.9% F1 on MSCOCO and 71.58% on AMBER.
title Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.07619