Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Sohee, Ryu, Soohyun, Park, Joonhyung, Yang, Eunho
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908519759872000
author Kim, Sohee
Ryu, Soohyun
Park, Joonhyung
Yang, Eunho
author_facet Kim, Sohee
Ryu, Soohyun
Park, Joonhyung
Yang, Eunho
contents Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evidence as being part of the image, leading to erroneous responses. In light of this finding, we probe whether LVLMs possess an internal capability to determine if textual concepts are grounded in the image, and discover a specific subset of Feed-Forward Network (FFN) neurons, termed Visual Absence-aware (VA) neurons, that consistently signal the visual absence through a distinctive activation pattern. Leveraging these patterns, we develop a detection module that systematically classifies whether an input token is visually grounded. Guided by its prediction, we propose a method to refine the outputs by reinterpreting question prompts or replacing the detected absent tokens during generation. Extensive experiments show that our method effectively mitigates the models' tendency to falsely presume the visual presence of text input and its generality across various LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
Kim, Sohee
Ryu, Soohyun
Park, Joonhyung
Yang, Eunho
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evidence as being part of the image, leading to erroneous responses. In light of this finding, we probe whether LVLMs possess an internal capability to determine if textual concepts are grounded in the image, and discover a specific subset of Feed-Forward Network (FFN) neurons, termed Visual Absence-aware (VA) neurons, that consistently signal the visual absence through a distinctive activation pattern. Leveraging these patterns, we develop a detection module that systematically classifies whether an input token is visually grounded. Guided by its prediction, we propose a method to refine the outputs by reinterpreting question prompts or replacing the detected absent tokens during generation. Extensive experiments show that our method effectively mitigates the models' tendency to falsely presume the visual presence of text input and its generality across various LVLMs.
title Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.03025