OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Mengdi, Qi, Zekun, Zhang, Shaochen, Zhang, Wenyao, Yu, Xinqiang, He, Jiawei, Wang, He, Yi, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917300114817024
author Jia, Mengdi
Qi, Zekun
Zhang, Shaochen
Zhang, Wenyao
Yu, Xinqiang
He, Jiawei
Wang, He
Yi, Li
author_facet Jia, Mengdi
Qi, Zekun
Zhang, Shaochen
Zhang, Wenyao
Yu, Xinqiang
He, Jiawei
Wang, He
Yi, Li
contents Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object counting, these tasks cover only the most elementary layer of spatial reasoning and are largely approaching saturation in the latest reasoning models. In this work, we introduce OmniSpatial, a comprehensive and challenging benchmark for spatial reasoning, grounded in cognitive psychology. OmniSpatial covers four major categories: dynamic reasoning, complex spatial logic, spatial interaction, and perspective-taking, with 50 fine-grained subcategories. Through careful manual annotation, we construct over 8.4K question-answer pairs. Extensive experiments show that both open- and closed-source VLMs exhibit significant limitations in comprehensive spatial reasoning. We also explore two strategies-PointGraph (explicit scene graph cues) and SpatialCoT (novel-view chain-of-thought)-to bolster spatial reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
Jia, Mengdi
Qi, Zekun
Zhang, Shaochen
Zhang, Wenyao
Yu, Xinqiang
He, Jiawei
Wang, He
Yi, Li
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object counting, these tasks cover only the most elementary layer of spatial reasoning and are largely approaching saturation in the latest reasoning models. In this work, we introduce OmniSpatial, a comprehensive and challenging benchmark for spatial reasoning, grounded in cognitive psychology. OmniSpatial covers four major categories: dynamic reasoning, complex spatial logic, spatial interaction, and perspective-taking, with 50 fine-grained subcategories. Through careful manual annotation, we construct over 8.4K question-answer pairs. Extensive experiments show that both open- and closed-source VLMs exhibit significant limitations in comprehensive spatial reasoning. We also explore two strategies-PointGraph (explicit scene graph cues) and SpatialCoT (novel-view chain-of-thought)-to bolster spatial reasoning.
title OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.03135