Investigating Spatial Attention Bias in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhary, Aryan, Goyal, Sanchit, Narang, Pratik, Kumar, Dhruv |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025)
by: Zhu, Yingjie, et al.
Published: (2025)
Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation
by: Gupta, Arnav, et al.
Published: (2025)
by: Gupta, Arnav, et al.
Published: (2025)
AffectSRNet : Facial Emotion-Aware Super-Resolution Network
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2025)
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2025)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026)
by: Ma, Chuang, et al.
Published: (2026)
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
by: Sasse, Kuleen, et al.
Published: (2024)
by: Sasse, Kuleen, et al.
Published: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
by: Stogiannidis, Ilias, et al.
Published: (2025)
by: Stogiannidis, Ilias, et al.
Published: (2025)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024)
by: Ray, Sourjyadip, et al.
Published: (2024)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
Exposing DeepFakes via Hyperspectral Domain Mapping
by: Mehta, Aditya, et al.
Published: (2025)
by: Mehta, Aditya, et al.
Published: (2025)
ColorFoil: Investigating Color Blindness in Large Vision and Language Models
by: Samin, Ahnaf Mozib, et al.
Published: (2024)
by: Samin, Ahnaf Mozib, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
Balancing the Scales: Enhancing Fairness in Facial Expression Recognition with Latent Alignment
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2024)
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2024)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
by: Wang, Zihu, et al.
Published: (2025)
by: Wang, Zihu, et al.
Published: (2025)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
by: Arif, Kazi Hasan Ibn, et al.
Published: (2025)
by: Arif, Kazi Hasan Ibn, et al.
Published: (2025)
Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
by: Goldshmidt, Roni
Published: (2025)
by: Goldshmidt, Roni
Published: (2025)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
AI Based Font Pair Suggestion Modelling For Graphic Design
by: Singh, Aryan, et al.
Published: (2025)
by: Singh, Aryan, et al.
Published: (2025)
Do Vision-Language Models Understand Compound Nouns?
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
by: Li, Hongxing, et al.
Published: (2025)
by: Li, Hongxing, et al.
Published: (2025)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
by: Yamada, Yutaro, et al.
Published: (2022)
by: Yamada, Yutaro, et al.
Published: (2022)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking
by: Li, Jingru, et al.
Published: (2026)
by: Li, Jingru, et al.
Published: (2026)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
by: Liu, Peiju, et al.
Published: (2026)
by: Liu, Peiju, et al.
Published: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
by: Jia, Mengdi, et al.
Published: (2025)
by: Jia, Mengdi, et al.
Published: (2025)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
by: Zhan, Shaoxiong, et al.
Published: (2026)
by: Zhan, Shaoxiong, et al.
Published: (2026)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
by: Fazli, Mehrdad, et al.
Published: (2025)
by: Fazli, Mehrdad, et al.
Published: (2025)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Similar Items
-
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025) -
Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation
by: Gupta, Arnav, et al.
Published: (2025) -
AffectSRNet : Facial Emotion-Aware Super-Resolution Network
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2025) -
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026) -
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
by: Xiong, Guangzhi, et al.
Published: (2026)