VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahgir, Haz Sameen, Chen, Xiaofu, Fu, Yu, Shayegani, Erfan, Abu-Ghazaleh, Nael, Kementchedjhieva, Yova, Dong, Yue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910130992316416
author Shahgir, Haz Sameen
Chen, Xiaofu
Fu, Yu
Shayegani, Erfan
Abu-Ghazaleh, Nael
Kementchedjhieva, Yova
Dong, Yue
author_facet Shahgir, Haz Sameen
Chen, Xiaofu
Fu, Yu
Shayegani, Erfan
Abu-Ghazaleh, Nael
Kementchedjhieva, Yova
Dong, Yue
contents Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information is still present in their internal representations. Prior work has attributed this ``hidden-in-plain-sight'' gap to the language model, but the cause remains unexplained. In this work, we demonstrate that this gap arises from the language model's lack of semantic labels for fine-grained visual details: when visual entities can be mapped to known concepts, VLMs bypass visual comparison and reason through language; when they cannot, VLMs resort to brittle and hallucinated descriptions. We verify this across semantic correspondence, synthetic shape matching, and face matching, and find that VLMs perform much better when the relevant entities are nameable than when they are unnamable. Mechanistically, Logit Lens analysis confirms that VLMs explicitly recover semantic labels for nameable entities and surface more unique tokens compared to unnameable entities. Furthermore, we show that this limitation can be addressed: teaching completely arbitrary names for unknown entities improves performance. More importantly, task-specific finetuning yields even stronger generalization without relying on language priors, i.e. through real visual perception. Our findings suggest that current VLM failures on visual tasks reflect a learned shortcut rather than a fundamental limitation of multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02486
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
Shahgir, Haz Sameen
Chen, Xiaofu
Fu, Yu
Shayegani, Erfan
Abu-Ghazaleh, Nael
Kementchedjhieva, Yova
Dong, Yue
Computer Vision and Pattern Recognition
Computation and Language
Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information is still present in their internal representations. Prior work has attributed this ``hidden-in-plain-sight'' gap to the language model, but the cause remains unexplained. In this work, we demonstrate that this gap arises from the language model's lack of semantic labels for fine-grained visual details: when visual entities can be mapped to known concepts, VLMs bypass visual comparison and reason through language; when they cannot, VLMs resort to brittle and hallucinated descriptions. We verify this across semantic correspondence, synthetic shape matching, and face matching, and find that VLMs perform much better when the relevant entities are nameable than when they are unnamable. Mechanistically, Logit Lens analysis confirms that VLMs explicitly recover semantic labels for nameable entities and surface more unique tokens compared to unnameable entities. Furthermore, we show that this limitation can be addressed: teaching completely arbitrary names for unknown entities improves performance. More importantly, task-specific finetuning yields even stronger generalization without relying on language priors, i.e. through real visual perception. Our findings suggest that current VLM failures on visual tasks reflect a learned shortcut rather than a fundamental limitation of multimodal reasoning.
title VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.02486