Large Vision-Language Models Get Lost in Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xi, Gongli, Tian, Ye, Yang, Mengyu, Yi, Huahui, Lin, Liang, Hao, Xiaoshuai, Wang, Kun, Wang, Wendong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910195969425408
author Xi, Gongli
Tian, Ye
Yang, Mengyu
Yi, Huahui
Lin, Liang
Hao, Xiaoshuai
Wang, Kun
Wang, Wendong
author_facet Xi, Gongli
Tian, Ye
Yang, Mengyu
Yi, Huahui
Lin, Liang
Hao, Xiaoshuai
Wang, Kun
Wang, Wendong
contents Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechanics and guiding architectural optimization. While prior statistical approaches have provided valuable attribution-based insights, they often lack a unified theoretical basis. To bridge this gap, we propose a unified framework grounded in information theory and geometry to quantify the geometric and entropic nature of residual updates. Applying this unified framework reveals a fundamental functional decoupling: Attention acts as a subspace-preserving operator focused on reconfiguration, whereas FFNs serve as subspace-expanding operators driving semantic innovation. Strikingly, further experiments demonstrate that replacing learned attention weights with predefined values (e.g., Gaussian noise) yields comparable or even superior performance across a majority of datasets relative to vanilla models. These results expose severe misallocation and redundancy in current mechanisms, suggesting that state-of-the-art LVLMs effectively ``get lost in attention'' rather than efficiently leveraging visual context.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05668
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large Vision-Language Models Get Lost in Attention
Xi, Gongli
Tian, Ye
Yang, Mengyu
Yi, Huahui
Lin, Liang
Hao, Xiaoshuai
Wang, Kun
Wang, Wendong
Artificial Intelligence
Computer Vision and Pattern Recognition
Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechanics and guiding architectural optimization. While prior statistical approaches have provided valuable attribution-based insights, they often lack a unified theoretical basis. To bridge this gap, we propose a unified framework grounded in information theory and geometry to quantify the geometric and entropic nature of residual updates. Applying this unified framework reveals a fundamental functional decoupling: Attention acts as a subspace-preserving operator focused on reconfiguration, whereas FFNs serve as subspace-expanding operators driving semantic innovation. Strikingly, further experiments demonstrate that replacing learned attention weights with predefined values (e.g., Gaussian noise) yields comparable or even superior performance across a majority of datasets relative to vanilla models. These results expose severe misallocation and redundancy in current mechanisms, suggesting that state-of-the-art LVLMs effectively ``get lost in attention'' rather than efficiently leveraging visual context.
title Large Vision-Language Models Get Lost in Attention
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.05668