SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wenyu, Ng, Wei En, Ma, Lixin, Wang, Yuwen, Zhao, Junqi, Koenecke, Allison, Li, Boyang, Wang, Lu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
by: Tiong, Anthony Meng Huat, et al.
Published: (2024)
by: Tiong, Anthony Meng Huat, et al.
Published: (2024)
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
by: Mao, Runhao, et al.
Published: (2026)
by: Mao, Runhao, et al.
Published: (2026)
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
On the Difficulty of Learning a Meta-network for Training Data Selection
by: Du, Zilin, et al.
Published: (2026)
by: Du, Zilin, et al.
Published: (2026)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
Hierarchical Spatial Proximity Reasoning for Vision-and-Language Navigation
by: Xu, Ming, et al.
Published: (2024)
by: Xu, Ming, et al.
Published: (2024)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
by: Lin, Junxiong, et al.
Published: (2024)
by: Lin, Junxiong, et al.
Published: (2024)
Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection
by: Lai, Runhe, et al.
Published: (2025)
by: Lai, Runhe, et al.
Published: (2025)
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
by: Liu, Zuyan, et al.
Published: (2024)
by: Liu, Zuyan, et al.
Published: (2024)
Spatiotemporal Blind-Spot Network with Calibrated Flow Alignment for Self-Supervised Video Denoising
by: Chen, Zikang, et al.
Published: (2024)
by: Chen, Zikang, et al.
Published: (2024)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
Blind Deconvolution for Color Images Using Normalized Quaternion Kernels
by: Yang, Yuming, et al.
Published: (2025)
by: Yang, Yuming, et al.
Published: (2025)
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
by: Huang, Xinmiao, et al.
Published: (2025)
by: Huang, Xinmiao, et al.
Published: (2025)
Sketch2MinSurf: Vision-Language Guided Generation of Editable Minimal Surfaces from Hand-Drawn Sketches
by: Wang, Wenda, et al.
Published: (2026)
by: Wang, Wenda, et al.
Published: (2026)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
Hierarchical GraphCut Phase Unwrapping based on Invariance of Diffeomorphisms Framework
by: Gao, Xiang, et al.
Published: (2025)
by: Gao, Xiang, et al.
Published: (2025)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
by: Wang, Jiayun, et al.
Published: (2026)
by: Wang, Jiayun, et al.
Published: (2026)
HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from Histopathology
by: Weng, Ziqiao, et al.
Published: (2025)
by: Weng, Ziqiao, et al.
Published: (2025)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
by: Jia, Hongrui, et al.
Published: (2026)
by: Jia, Hongrui, et al.
Published: (2026)
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Into the Unknown: Accounting for Missing Demographic Data when Mitigating Ad Delivery Skew
by: Corpus, Isabel, et al.
Published: (2026)
by: Corpus, Isabel, et al.
Published: (2026)
Deep Generative Models Unveil Patterns in Medical Images Through Vision-Language Conditioning
by: Xing, Xiaodan, et al.
Published: (2024)
by: Xing, Xiaodan, et al.
Published: (2024)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
by: Wang, Anmin, et al.
Published: (2026)
by: Wang, Anmin, et al.
Published: (2026)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
by: Wei, Xingguang, et al.
Published: (2025)
by: Wei, Xingguang, et al.
Published: (2025)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
by: Khangaonkar, Om, et al.
Published: (2026)
by: Khangaonkar, Om, et al.
Published: (2026)
Vision-Language Memory for Spatial Reasoning
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium
by: Xiao, Qingxin, et al.
Published: (2026)
by: Xiao, Qingxin, et al.
Published: (2026)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
by: Lin, Junxiong, et al.
Published: (2024)
by: Lin, Junxiong, et al.
Published: (2024)
Similar Items
-
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026) -
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
by: Tiong, Anthony Meng Huat, et al.
Published: (2024) -
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024) -
The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
by: Mao, Runhao, et al.
Published: (2026) -
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025)