Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shiqi, Zhu, Tongyao, Zhou, Ruochen, Zhang, Jinghan, Gao, Siyang, Niebles, Juan Carlos, Geva, Mor, He, Junxian, Wu, Jiajun, Li, Manling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026)
by: Chen, Shiqi, et al.
Published: (2026)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
by: Chen, Shiqi, et al.
Published: (2024)
by: Chen, Shiqi, et al.
Published: (2024)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Universal Jailbreak Suffixes Are Strong Attention Hijackers
by: Ben-Tov, Matan, et al.
Published: (2025)
by: Ben-Tov, Matan, et al.
Published: (2025)
Estimating Knowledge in Large Language Models Without Generating a Single Token
by: Gottesman, Daniela, et al.
Published: (2024)
by: Gottesman, Daniela, et al.
Published: (2024)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Language Models Encode Numbers Using Digit Representations in Base 10
by: Levy, Amit Arnold, et al.
Published: (2024)
by: Levy, Amit Arnold, et al.
Published: (2024)
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
by: Zhou, Ruochen, et al.
Published: (2025)
by: Zhou, Ruochen, et al.
Published: (2025)
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026)
by: Rai, Daking, et al.
Published: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
by: Ramati, Dana, et al.
Published: (2024)
by: Ramati, Dana, et al.
Published: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
by: Barbi, Ohav, et al.
Published: (2025)
by: Barbi, Ohav, et al.
Published: (2025)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
by: Yang, Sohee, et al.
Published: (2024)
by: Yang, Sohee, et al.
Published: (2024)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
by: Yang, Sohee, et al.
Published: (2024)
by: Yang, Sohee, et al.
Published: (2024)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
by: Grosbard, Idan Daniel, et al.
Published: (2026)
by: Grosbard, Idan Daniel, et al.
Published: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
Detecting (Un)answerability in Large Language Models with Linear Directions
by: Lavi, Maor Juliet, et al.
Published: (2025)
by: Lavi, Maor Juliet, et al.
Published: (2025)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2026)
by: Gekhman, Zorik, et al.
Published: (2026)
Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts
by: Zhu, Youxiang, et al.
Published: (2025)
by: Zhu, Youxiang, et al.
Published: (2025)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
by: Xue, Qiyao, et al.
Published: (2025)
by: Xue, Qiyao, et al.
Published: (2025)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024)
by: Zhou, Qiji, et al.
Published: (2024)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
by: Cohen, Ido, et al.
Published: (2024)
by: Cohen, Ido, et al.
Published: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Similar Items
-
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
by: Chen, Shiqi, et al.
Published: (2025) -
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025) -
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026) -
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026) -
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)