The Spatial Blindspot of Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alam, Nahid, Murali, Leema Krishna, Bharadwaj, Siddhant, Liu, Patrick, Chung, Timothy, Sharma, Drishti, A, Akshata, Kiran, Kranthi, Tam, Wesley, Vegesna, Bala Krishna S |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
Behind Maya: Building a Multilingual Vision Language Model
von: Alam, Nahid, et al.
Veröffentlicht: (2025)
von: Alam, Nahid, et al.
Veröffentlicht: (2025)
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
Maya: An Instruction Finetuned Multilingual Multimodal Model
von: Alam, Nahid, et al.
Veröffentlicht: (2024)
von: Alam, Nahid, et al.
Veröffentlicht: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
Adaptive Multi Scale Document Binarisation Using Vision Mamba
von: Azfar, Mohd., et al.
Veröffentlicht: (2024)
von: Azfar, Mohd., et al.
Veröffentlicht: (2024)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
Causal Physics Steering in Video World Models via Concept Activation Vectors
von: Alam, Nahid
Veröffentlicht: (2026)
von: Alam, Nahid
Veröffentlicht: (2026)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping
von: Vegesna, Nitin, et al.
Veröffentlicht: (2026)
von: Vegesna, Nitin, et al.
Veröffentlicht: (2026)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
Embedding Geometries of Contrastive Language-Image Pre-Training
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
Enhancing Vision Language Models with Logic Reasoning for Situational Awareness
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
The Nonverbal Gap: Toward Affective Computer Vision for Safer and More Equitable Online Dating
von: Kandala, Ratna, et al.
Veröffentlicht: (2026)
von: Kandala, Ratna, et al.
Veröffentlicht: (2026)
PRECISe : Prototype-Reservation for Explainable Classification under Imbalanced and Scarce-Data Settings
von: Ganatra, Vaibhav, et al.
Veröffentlicht: (2024)
von: Ganatra, Vaibhav, et al.
Veröffentlicht: (2024)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
von: Lim, Gyubeum, et al.
Veröffentlicht: (2025)
von: Lim, Gyubeum, et al.
Veröffentlicht: (2025)
From Steering to Pedalling: Do Autonomous Driving VLMs Generalize to Cyclist-Assistive Spatial Perception and Planning?
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2026)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2026)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
Accelerating HEVC Intra Partitioning via a CNN-Hierarchical Attention Transformer Hybrid
von: Sharma, Krishna Kumar, et al.
Veröffentlicht: (2026)
von: Sharma, Krishna Kumar, et al.
Veröffentlicht: (2026)
Watermarking in Diffusion Model: Gaussian Shading with Exact Diffusion Inversion via Coupled Transformations (EDICT)
von: Panthi, Krishna
Veröffentlicht: (2025)
von: Panthi, Krishna
Veröffentlicht: (2025)
Grounded 3D-Aware Spatial Vision-Language Modeling
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
Dynamic Contextual Attention Network: Transforming Spatial Representations into Adaptive Insights for Endoscopic Polyp Diagnosis
von: Cherukuri, Teja Krishna, et al.
Veröffentlicht: (2025)
von: Cherukuri, Teja Krishna, et al.
Veröffentlicht: (2025)
Fill in the ____ (a Diffusion-based Image Inpainting Pipeline)
von: Gebre, Eyoel, et al.
Veröffentlicht: (2024)
von: Gebre, Eyoel, et al.
Veröffentlicht: (2024)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
SPACE: 3D Spatial Co-operation and Exploration Framework for Robust Mapping and Coverage with Multi-Robot Systems
von: Ghanta, Sai Krishna, et al.
Veröffentlicht: (2024)
von: Ghanta, Sai Krishna, et al.
Veröffentlicht: (2024)
MRI Volume-Based Robust Brain Age Estimation Using Weight-Shared Spatial Attention in 3D CNNs
von: Kancharla, Vamshi Krishna, et al.
Veröffentlicht: (2024)
von: Kancharla, Vamshi Krishna, et al.
Veröffentlicht: (2024)
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
von: Nakka, Krishna Kanth
Veröffentlicht: (2025)
von: Nakka, Krishna Kanth
Veröffentlicht: (2025)
YoChameleon: Personalized Vision and Language Generation
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
CanViT: Toward Active-Vision Foundation Models
von: Berreby, Yohaï-Eliel, et al.
Veröffentlicht: (2026)
von: Berreby, Yohaï-Eliel, et al.
Veröffentlicht: (2026)
Optimizing Latent Graph Representations of Surgical Scenes for Zero-Shot Domain Transfer
von: Satyanaik, Siddhant, et al.
Veröffentlicht: (2024)
von: Satyanaik, Siddhant, et al.
Veröffentlicht: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
von: Su, Xia, et al.
Veröffentlicht: (2026)
von: Su, Xia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
von: Alam, Nahid, et al.
Veröffentlicht: (2026) -
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
von: Zafar, Anas, et al.
Veröffentlicht: (2026) -
Behind Maya: Building a Multilingual Vision Language Model
von: Alam, Nahid, et al.
Veröffentlicht: (2025) -
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025) -
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)