How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912700691382272 |
|---|---|
| author | Yu, Songsong Chen, Yuxin Ju, Hao Jia, Lianjie Zhang, Fuxi Huang, Shaofei Wu, Yuhan Cui, Rundi Ran, Binghao Zhang, Zaibin Zheng, Zhedong Zhang, Zhipeng Wang, Yifan Song, Lin Wang, Lijun Li, Yanwei Shan, Ying Lu, Huchuan |
| author_facet | Yu, Songsong Chen, Yuxin Ju, Hao Jia, Lianjie Zhang, Fuxi Huang, Shaofei Wu, Yuhan Cui, Rundi Ran, Binghao Zhang, Zaibin Zheng, Zhedong Zhang, Zhipeng Wang, Yifan Song, Lin Wang, Lijun Li, Yanwei Shan, Ying Lu, Huchuan |
| contents | Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR remains highly challenging due to the complexity of representing and reasoning over three-dimensional space. In this paper, we present a systematic investigation of VSR in VLMs, encompassing a review of existing methodologies across input modalities, model architectures, training strategies, and reasoning mechanisms. Furthermore, we categorize spatial intelligence into three levels of capability, ie, basic perception, spatial understanding, spatial planning, and curate SIBench, a spatial intelligence benchmark encompassing nearly 20 open-source datasets across 23 task settings. Experiments with state-of-the-art VLMs reveal a pronounced gap between perception and reasoning, as models show competence in basic perceptual tasks but consistently underperform in understanding and planning tasks, particularly in numerical estimation, multi-view reasoning, temporal dynamics, and spatial imagination. These findings underscore the substantial challenges that remain in achieving spatial intelligence, while providing both a systematic roadmap and a comprehensive benchmark to drive future research in the field. The related resources of this study are accessible at https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18905 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective Yu, Songsong Chen, Yuxin Ju, Hao Jia, Lianjie Zhang, Fuxi Huang, Shaofei Wu, Yuhan Cui, Rundi Ran, Binghao Zhang, Zaibin Zheng, Zhedong Zhang, Zhipeng Wang, Yifan Song, Lin Wang, Lijun Li, Yanwei Shan, Ying Lu, Huchuan Artificial Intelligence Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR remains highly challenging due to the complexity of representing and reasoning over three-dimensional space. In this paper, we present a systematic investigation of VSR in VLMs, encompassing a review of existing methodologies across input modalities, model architectures, training strategies, and reasoning mechanisms. Furthermore, we categorize spatial intelligence into three levels of capability, ie, basic perception, spatial understanding, spatial planning, and curate SIBench, a spatial intelligence benchmark encompassing nearly 20 open-source datasets across 23 task settings. Experiments with state-of-the-art VLMs reveal a pronounced gap between perception and reasoning, as models show competence in basic perceptual tasks but consistently underperform in understanding and planning tasks, particularly in numerical estimation, multi-view reasoning, temporal dynamics, and spatial imagination. These findings underscore the substantial challenges that remain in achieving spatial intelligence, while providing both a systematic roadmap and a comprehensive benchmark to drive future research in the field. The related resources of this study are accessible at https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/. |
| title | How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2509.18905 |