Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Yingjie, Bai, Xuefeng, Chen, Kehai, Xiang, Yang, Pan, Youcheng, Hou, Yongshuai, Guan, Weili, Yu, Jun, Zhang, Min
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918320306913280
author Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Pan, Youcheng
Hou, Yongshuai
Guan, Weili
Yu, Jun
Zhang, Min
author_facet Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Pan, Youcheng
Hou, Yongshuai
Guan, Weili
Yu, Jun
Zhang, Min
contents Large Vision-Language Models (LVLMs) have achieved remarkable success across a wide range of multimodal tasks, yet their robustness to spatial variations remains insufficiently understood. In this work, we conduct a systematic study of the spatial bias of LVLMs, examining how models respond when identical key visual information is placed at different locations within an image. Through controlled probing experiments, we observe that current LVLMs often produce inconsistent outputs under such spatial shifts, revealing a clear spatial bias in their semantic understanding. Further analysis indicates that this bias does not stem from the vision encoder, but rather from a mismatch in attention mechanisms between the vision encoder and the large language model, which disrupts the global information flow. Motivated by this insight, we propose Adaptive Global Context Injection (AGCI), a lightweight mechanism that dynamically injects shared global visual context into each image token. AGCI works without architectural modifications, mitigating spatial bias by enhancing the semantic accessibility of image tokens while preserving the model's intrinsic capabilities. Extensive experiments demonstrate that AGCI not only enhances the spatial robustness of LVLMs, but also achieves strong performance on various downstream tasks and hallucination benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Pan, Youcheng
Hou, Yongshuai
Guan, Weili
Yu, Jun
Zhang, Min
Computer Vision and Pattern Recognition
Computation and Language
Large Vision-Language Models (LVLMs) have achieved remarkable success across a wide range of multimodal tasks, yet their robustness to spatial variations remains insufficiently understood. In this work, we conduct a systematic study of the spatial bias of LVLMs, examining how models respond when identical key visual information is placed at different locations within an image. Through controlled probing experiments, we observe that current LVLMs often produce inconsistent outputs under such spatial shifts, revealing a clear spatial bias in their semantic understanding. Further analysis indicates that this bias does not stem from the vision encoder, but rather from a mismatch in attention mechanisms between the vision encoder and the large language model, which disrupts the global information flow. Motivated by this insight, we propose Adaptive Global Context Injection (AGCI), a lightweight mechanism that dynamically injects shared global visual context into each image token. AGCI works without architectural modifications, mitigating spatial bias by enhancing the semantic accessibility of image tokens while preserving the model's intrinsic capabilities. Extensive experiments demonstrate that AGCI not only enhances the spatial robustness of LVLMs, but also achieves strong performance on various downstream tasks and hallucination benchmarks.
title Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.21984