How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Songsong, Chen, Yuxin, Ju, Hao, Jia, Lianjie, Zhang, Fuxi, Huang, Shaofei, Wu, Yuhan, Cui, Rundi, Ran, Binghao, Zhang, Zaibin, Zheng, Zhedong, Zhang, Zhipeng, Wang, Yifan, Song, Lin, Wang, Lijun, Li, Yanwei, Shan, Ying, Lu, Huchuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912700691382272
author Yu, Songsong
Chen, Yuxin
Ju, Hao
Jia, Lianjie
Zhang, Fuxi
Huang, Shaofei
Wu, Yuhan
Cui, Rundi
Ran, Binghao
Zhang, Zaibin
Zheng, Zhedong
Zhang, Zhipeng
Wang, Yifan
Song, Lin
Wang, Lijun
Li, Yanwei
Shan, Ying
Lu, Huchuan
author_facet Yu, Songsong
Chen, Yuxin
Ju, Hao
Jia, Lianjie
Zhang, Fuxi
Huang, Shaofei
Wu, Yuhan
Cui, Rundi
Ran, Binghao
Zhang, Zaibin
Zheng, Zhedong
Zhang, Zhipeng
Wang, Yifan
Song, Lin
Wang, Lijun
Li, Yanwei
Shan, Ying
Lu, Huchuan
contents Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR remains highly challenging due to the complexity of representing and reasoning over three-dimensional space. In this paper, we present a systematic investigation of VSR in VLMs, encompassing a review of existing methodologies across input modalities, model architectures, training strategies, and reasoning mechanisms. Furthermore, we categorize spatial intelligence into three levels of capability, ie, basic perception, spatial understanding, spatial planning, and curate SIBench, a spatial intelligence benchmark encompassing nearly 20 open-source datasets across 23 task settings. Experiments with state-of-the-art VLMs reveal a pronounced gap between perception and reasoning, as models show competence in basic perceptual tasks but consistently underperform in understanding and planning tasks, particularly in numerical estimation, multi-view reasoning, temporal dynamics, and spatial imagination. These findings underscore the substantial challenges that remain in achieving spatial intelligence, while providing both a systematic roadmap and a comprehensive benchmark to drive future research in the field. The related resources of this study are accessible at https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Yu, Songsong
Chen, Yuxin
Ju, Hao
Jia, Lianjie
Zhang, Fuxi
Huang, Shaofei
Wu, Yuhan
Cui, Rundi
Ran, Binghao
Zhang, Zaibin
Zheng, Zhedong
Zhang, Zhipeng
Wang, Yifan
Song, Lin
Wang, Lijun
Li, Yanwei
Shan, Ying
Lu, Huchuan
Artificial Intelligence
Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR remains highly challenging due to the complexity of representing and reasoning over three-dimensional space. In this paper, we present a systematic investigation of VSR in VLMs, encompassing a review of existing methodologies across input modalities, model architectures, training strategies, and reasoning mechanisms. Furthermore, we categorize spatial intelligence into three levels of capability, ie, basic perception, spatial understanding, spatial planning, and curate SIBench, a spatial intelligence benchmark encompassing nearly 20 open-source datasets across 23 task settings. Experiments with state-of-the-art VLMs reveal a pronounced gap between perception and reasoning, as models show competence in basic perceptual tasks but consistently underperform in understanding and planning tasks, particularly in numerical estimation, multi-view reasoning, temporal dynamics, and spatial imagination. These findings underscore the substantial challenges that remain in achieving spatial intelligence, while providing both a systematic roadmap and a comprehensive benchmark to drive future research in the field. The related resources of this study are accessible at https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/.
title How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
topic Artificial Intelligence
url https://arxiv.org/abs/2509.18905