HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Huizhi, Shen, Yichao, Deng, Yu, Xu, Sicheng, Feng, Zhiyuan, Zhang, Tong, Liang, Yaobo, Yang, Jiaolong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911546235420672
author Liang, Huizhi
Shen, Yichao
Deng, Yu
Xu, Sicheng
Feng, Zhiyuan
Zhang, Tong
Liang, Yaobo
Yang, Jiaolong
author_facet Liang, Huizhi
Shen, Yichao
Deng, Yu
Xu, Sicheng
Feng, Zhiyuan
Zhang, Tong
Liang, Yaobo
Yang, Jiaolong
contents Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that decomposes the learning of 3D spatial understanding in VLMs into four progressively complex levels, from geometric perception to abstract spatial reasoning. Guided by this framework, we construct an automated pipeline that processes approximately 5M images with over 45M objects to generate 3D spatial VQA pairs across diverse tasks and scenes for VLM supervised fine-tuning. We also develop an RGB-D VLM incorporating metric-scale point maps as auxiliary inputs to further enhance spatial understanding. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on multiple spatial understanding and reasoning benchmarks, surpassing specialized spatial models and large proprietary systems such as Gemini-2.5-pro and GPT-5. Moreover, our analysis reveals clear dependencies among hierarchical task levels, offering new insights into how multi-level task design facilitates the emergence of 3D spatial intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25411
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
Liang, Huizhi
Shen, Yichao
Deng, Yu
Xu, Sicheng
Feng, Zhiyuan
Zhang, Tong
Liang, Yaobo
Yang, Jiaolong
Computer Vision and Pattern Recognition
Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that decomposes the learning of 3D spatial understanding in VLMs into four progressively complex levels, from geometric perception to abstract spatial reasoning. Guided by this framework, we construct an automated pipeline that processes approximately 5M images with over 45M objects to generate 3D spatial VQA pairs across diverse tasks and scenes for VLM supervised fine-tuning. We also develop an RGB-D VLM incorporating metric-scale point maps as auxiliary inputs to further enhance spatial understanding. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on multiple spatial understanding and reasoning benchmarks, surpassing specialized spatial models and large proprietary systems such as Gemini-2.5-pro and GPT-5. Moreover, our analysis reveals clear dependencies among hierarchical task levels, offering new insights into how multi-level task design facilitates the emergence of 3D spatial intelligence.
title HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25411