Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Shiqi, Zhu, Tongyao, Zhou, Ruochen, Zhang, Jinghan, Gao, Siyang, Niebles, Juan Carlos, Geva, Mor, He, Junxian, Wu, Jiajun, Li, Manling
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909841859018752
author Chen, Shiqi
Zhu, Tongyao
Zhou, Ruochen
Zhang, Jinghan
Gao, Siyang
Niebles, Juan Carlos
Geva, Mor
He, Junxian
Wu, Jiajun
Li, Manling
author_facet Chen, Shiqi
Zhu, Tongyao
Zhou, Ruochen
Zhang, Jinghan
Gao, Siyang
Niebles, Juan Carlos
Geva, Mor
He, Junxian
Wu, Jiajun
Li, Manling
contents Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant challenges for current VLMs. In this work, we study the spatial reasoning challenge from the lens of mechanistic interpretability, diving into the model's internal states to examine the interactions between image and text tokens. By tracing attention distribution over the image through out intermediate layers, we observe that successful spatial reasoning correlates strongly with the model's ability to align its attention distribution with actual object locations, particularly differing between familiar and unfamiliar spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when confident, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible cost. We make code and data publicly available for research purposes at https://github.com/shiqichen17/AdaptVis.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
Chen, Shiqi
Zhu, Tongyao
Zhou, Ruochen
Zhang, Jinghan
Gao, Siyang
Niebles, Juan Carlos
Geva, Mor
He, Junxian
Wu, Jiajun
Li, Manling
Computation and Language
Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant challenges for current VLMs. In this work, we study the spatial reasoning challenge from the lens of mechanistic interpretability, diving into the model's internal states to examine the interactions between image and text tokens. By tracing attention distribution over the image through out intermediate layers, we observe that successful spatial reasoning correlates strongly with the model's ability to align its attention distribution with actual object locations, particularly differing between familiar and unfamiliar spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when confident, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible cost. We make code and data publicly available for research purposes at https://github.com/shiqichen17/AdaptVis.
title Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
topic Computation and Language
url https://arxiv.org/abs/2503.01773