Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xueqi, Yang, Shuo, Jiang, Yanbei, Liu, Shu, Liu, Zhenzhen, Ao, Jiayang, Ma, Xingjun, Erfani, Sarah Monazam, Bailey, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Unlearnable Examples For Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
End-to-End Anti-Backdoor Learning on Images and Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
LDReg: Local Dimensionality Regularized Self-Supervised Learning
by: Huang, Hanxun, et al.
Published: (2024)
by: Huang, Hanxun, et al.
Published: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Efficient Neural Implicit Representation for 3D Human Reconstruction
by: Huang, Zexu, et al.
Published: (2024)
by: Huang, Zexu, et al.
Published: (2024)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
AudioMosaic: Contrastive Masked Audio Representation Learning
by: Huang, Hanxun, et al.
Published: (2026)
by: Huang, Hanxun, et al.
Published: (2026)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
Exploring Weak-to-Strong Generalization for CLIP-based Classification
by: Li, Jinhao, et al.
Published: (2025)
by: Li, Jinhao, et al.
Published: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
The Role of Quantum Measurements when Testing the Quantum Nature of Gravity
by: Miki, Daisuke, et al.
Published: (2025)
by: Miki, Daisuke, et al.
Published: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
AIR: Post-training Data Selection for Reasoning via Attention Head Influence
by: Liu, Jinrui, et al.
Published: (2025)
by: Liu, Jinrui, et al.
Published: (2025)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
by: Mou, Tingshu, et al.
Published: (2026)
by: Mou, Tingshu, et al.
Published: (2026)
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
by: Nam, Andrew, et al.
Published: (2025)
by: Nam, Andrew, et al.
Published: (2025)
SafeCoT: Improving VLM Safety with Minimal Reasoning
by: Ma, Jiachen, et al.
Published: (2025)
by: Ma, Jiachen, et al.
Published: (2025)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
When and Why Grouping Attention Heads Accelerates Muon Optimization
by: Zhang, Hongtao, et al.
Published: (2026)
by: Zhang, Hongtao, et al.
Published: (2026)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
by: Liu, Shuhang, et al.
Published: (2025)
by: Liu, Shuhang, et al.
Published: (2025)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving
by: Liu, Haichao, et al.
Published: (2025)
by: Liu, Haichao, et al.
Published: (2025)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Testing the quantum nature of gravity through interferometry
by: Liu, Yubao, et al.
Published: (2025)
by: Liu, Yubao, et al.
Published: (2025)
Semiclassical gravity phenomenology under the causal-conditional quantum measurement prescription II: Heisenberg picture and apparent optical entanglement
by: Liu, Yubao, et al.
Published: (2024)
by: Liu, Yubao, et al.
Published: (2024)
Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
by: Fan, Hongxing, et al.
Published: (2025)
by: Fan, Hongxing, et al.
Published: (2025)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion
by: Wang, Ruofan, et al.
Published: (2025)
by: Wang, Ruofan, et al.
Published: (2025)
From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning
by: Xu, Zonghuan, et al.
Published: (2026)
by: Xu, Zonghuan, et al.
Published: (2026)
Similar Items
-
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
by: Ma, Xueqi, et al.
Published: (2025) -
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
by: Jiang, Yanbei, et al.
Published: (2025) -
Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
by: Ma, Xueqi, et al.
Published: (2025) -
Unlearnable Examples For Time Series
by: Jiang, Yujing, et al.
Published: (2024) -
End-to-End Anti-Backdoor Learning on Images and Time Series
by: Jiang, Yujing, et al.
Published: (2024)