How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Songsong, Chen, Yuxin, Ju, Hao, Jia, Lianjie, Zhang, Fuxi, Huang, Shaofei, Wu, Yuhan, Cui, Rundi, Ran, Binghao, Zhang, Zaibin, Zheng, Zhedong, Zhang, Zhipeng, Wang, Yifan, Song, Lin, Wang, Lijun, Li, Yanwei, Shan, Ying, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think3D: Thinking with Space for Spatial Reasoning
by: Zhang, Zaibin, et al.
Published: (2026)
by: Zhang, Zaibin, et al.
Published: (2026)
AR-MOT: Autoregressive Multi-object Tracking
by: Jia, Lianjie, et al.
Published: (2026)
by: Jia, Lianjie, et al.
Published: (2026)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
by: Zhang, Zaibin, et al.
Published: (2023)
by: Zhang, Zaibin, et al.
Published: (2023)
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
by: Zhang, Zaibin, et al.
Published: (2024)
by: Zhang, Zaibin, et al.
Published: (2024)
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
by: Yu, Songsong, et al.
Published: (2025)
by: Yu, Songsong, et al.
Published: (2025)
Revisiting Salient Object Detection from an Observer-Centric Perspective
by: Zhang, Fuxi, et al.
Published: (2026)
by: Zhang, Fuxi, et al.
Published: (2026)
Semantic Generative Tuning for Unified Multimodal Models
by: Yu, Songsong, et al.
Published: (2026)
by: Yu, Songsong, et al.
Published: (2026)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
by: Zhang, Zaibin, et al.
Published: (2024)
by: Zhang, Zaibin, et al.
Published: (2024)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
by: Ju, Hao, et al.
Published: (2024)
by: Ju, Hao, et al.
Published: (2024)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
by: Ju, Hao, et al.
Published: (2025)
by: Ju, Hao, et al.
Published: (2025)
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
by: Zhao, Zhida, et al.
Published: (2025)
by: Zhao, Zhida, et al.
Published: (2025)
Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
From Instruction to Event: Sound-Triggered Mobile Manipulation
by: Ju, Hao, et al.
Published: (2026)
by: Ju, Hao, et al.
Published: (2026)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
by: Zhang, Yusen, et al.
Published: (2025)
by: Zhang, Yusen, et al.
Published: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
by: Zhao, Hengrun, et al.
Published: (2025)
by: Zhao, Hengrun, et al.
Published: (2025)
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
3D Primitives are a Spatial Language for VLMs
by: Liu, Junze, et al.
Published: (2026)
by: Liu, Junze, et al.
Published: (2026)
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
by: Wang, Junran, et al.
Published: (2026)
by: Wang, Junran, et al.
Published: (2026)
Spatial Semantic Recurrent Mining for Referring Image Segmentation
by: Yang, Jiaxing, et al.
Published: (2024)
by: Yang, Jiaxing, et al.
Published: (2024)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
by: Zhang, Ruiyang, et al.
Published: (2024)
by: Zhang, Ruiyang, et al.
Published: (2024)
A Spatial Calibration Method for Robust Cooperative Perception
by: Song, Zhiying, et al.
Published: (2023)
by: Song, Zhiying, et al.
Published: (2023)
Bayesian Calibration of the Intelligent Driver Model
by: Zhang, Chengyuan, et al.
Published: (2022)
by: Zhang, Chengyuan, et al.
Published: (2022)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
by: Li, Jiangyang, et al.
Published: (2026)
by: Li, Jiangyang, et al.
Published: (2026)
Assessment of Multimodal Large Language Models in Alignment with Human Values
by: Shi, Zhelun, et al.
Published: (2024)
by: Shi, Zhelun, et al.
Published: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Unlocking Dense Metric Depth Estimation in VLMs
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
Lightweight RGB-D Salient Object Detection from a Speed-Accuracy Tradeoff Perspective
by: Duan, Songsong, et al.
Published: (2025)
by: Duan, Songsong, et al.
Published: (2025)
Similar Items
-
Think3D: Thinking with Space for Spatial Reasoning
by: Zhang, Zaibin, et al.
Published: (2026) -
AR-MOT: Autoregressive Multi-object Tracking
by: Jia, Lianjie, et al.
Published: (2026) -
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
by: Zhang, Zaibin, et al.
Published: (2023) -
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
by: Zhang, Zaibin, et al.
Published: (2024) -
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
by: Yu, Songsong, et al.
Published: (2025)