ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Dingming, Li, Hongxing, Wang, Zixuan, Yan, Yuchen, Zhang, Hang, Chen, Siqi, Hou, Guiyang, Jiang, Shengpei, Zhang, Wenqi, Shen, Yongliang, Lu, Weiming, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916978222956544
author Li, Dingming
Li, Hongxing
Wang, Zixuan
Yan, Yuchen
Zhang, Hang
Chen, Siqi
Hou, Guiyang
Jiang, Shengpei
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Zhuang, Yueting
author_facet Li, Dingming
Li, Hongxing
Wang, Zixuan
Yan, Yuchen
Zhang, Hang
Chen, Siqi
Hou, Guiyang
Jiang, Shengpei
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Zhuang, Yueting
contents Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We identify a critical limitation: current VLMs excel primarily at egocentric spatial reasoning (from the camera's perspective) but fail to generalize to allocentric viewpoints when required to adopt another entity's spatial frame of reference. We introduce ViewSpatial-Bench, the first comprehensive benchmark designed specifically for multi-viewpoint spatial localization recognition evaluation across five distinct task types, supported by an automated 3D annotation pipeline that generates precise directional labels. Comprehensive evaluation of diverse VLMs on ViewSpatial-Bench reveals a significant performance disparity: models demonstrate reasonable performance on camera-perspective tasks but exhibit reduced accuracy when reasoning from a human viewpoint. By fine-tuning VLMs on our multi-perspective spatial dataset, we achieve an overall performance improvement of 46.24% across tasks, highlighting the efficacy of our approach. Our work establishes a crucial benchmark for spatial intelligence in embodied AI systems and provides empirical evidence that modeling 3D spatial relationships enhances VLMs' corresponding spatial comprehension capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Li, Dingming
Li, Hongxing
Wang, Zixuan
Yan, Yuchen
Zhang, Hang
Chen, Siqi
Hou, Guiyang
Jiang, Shengpei
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Zhuang, Yueting
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We identify a critical limitation: current VLMs excel primarily at egocentric spatial reasoning (from the camera's perspective) but fail to generalize to allocentric viewpoints when required to adopt another entity's spatial frame of reference. We introduce ViewSpatial-Bench, the first comprehensive benchmark designed specifically for multi-viewpoint spatial localization recognition evaluation across five distinct task types, supported by an automated 3D annotation pipeline that generates precise directional labels. Comprehensive evaluation of diverse VLMs on ViewSpatial-Bench reveals a significant performance disparity: models demonstrate reasonable performance on camera-perspective tasks but exhibit reduced accuracy when reasoning from a human viewpoint. By fine-tuning VLMs on our multi-perspective spatial dataset, we achieve an overall performance improvement of 46.24% across tasks, highlighting the efficacy of our approach. Our work establishes a crucial benchmark for spatial intelligence in embodied AI systems and provides empirical evidence that modeling 3D spatial relationships enhances VLMs' corresponding spatial comprehension capabilities.
title ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.21500