SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Peiwen, Lang, Shiqiang, Wu, Dongming, Ding, Yi, Feng, Kaituo, Liu, Huadai, Ye, Zhen, Liu, Rui, Liu, Yun-Hui, Wang, Jianan, Yue, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916045900480512
author Sun, Peiwen
Lang, Shiqiang
Wu, Dongming
Ding, Yi
Feng, Kaituo
Liu, Huadai
Ye, Zhen
Liu, Rui
Liu, Yun-Hui
Wang, Jianan
Yue, Xiangyu
author_facet Sun, Peiwen
Lang, Shiqiang
Wu, Dongming
Ding, Yi
Feng, Kaituo
Liu, Huadai
Ye, Zhen
Liu, Rui
Liu, Yun-Hui
Wang, Jianan
Yue, Xiangyu
contents With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to advance all-scale spatial reasoning across diverse scenarios by tackling two key challenges: 1) the heavy reliance on indoor 3D scans and labor-intensive manual annotations for dataset curation; 2) the absence of effective all-scale scene modeling, which often leads to overfitting to individual scenes. In this paper, we introduce a holistic solution that integrates a structured spatial reasoning knowledge system, scale-aware modeling, and a progressive training paradigm, as the first attempt to broaden the all-scale spatial intelligence of MLLMs to the best of our knowledge. Using a task-specific, specialist-driven automated pipeline, we curate over 38K video scenes across 5 spatial scales to create SpaceVista-1M, a dataset comprising approximately 1M spatial QA pairs spanning 19 diverse task types. While specialist models can inject useful domain knowledge, they are not reliable for evaluation. We then build an all-scale benchmark with precise annotations by manually recording, retrieving, and assembling video-based data. However, naive training with SpaceVista-1M often yields suboptimal results due to the potential knowledge conflict. Accordingly, we introduce SpaceVista-7B, a spatial reasoning model that accepts dense inputs beyond semantics and uses scale as an anchor for scale-aware experts and progressive rewards. Finally, extensive evaluations across 5 benchmarks, including our SpaceVista-Bench, demonstrate competitive performance, showcasing strong generalization across all scales and scenarios. Our dataset, model, and benchmark will be released on https://peiwensun2000.github.io/mm2km .
format Preprint
id arxiv_https___arxiv_org_abs_2510_09606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
Sun, Peiwen
Lang, Shiqiang
Wu, Dongming
Ding, Yi
Feng, Kaituo
Liu, Huadai
Ye, Zhen
Liu, Rui
Liu, Yun-Hui
Wang, Jianan
Yue, Xiangyu
Computer Vision and Pattern Recognition
With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to advance all-scale spatial reasoning across diverse scenarios by tackling two key challenges: 1) the heavy reliance on indoor 3D scans and labor-intensive manual annotations for dataset curation; 2) the absence of effective all-scale scene modeling, which often leads to overfitting to individual scenes. In this paper, we introduce a holistic solution that integrates a structured spatial reasoning knowledge system, scale-aware modeling, and a progressive training paradigm, as the first attempt to broaden the all-scale spatial intelligence of MLLMs to the best of our knowledge. Using a task-specific, specialist-driven automated pipeline, we curate over 38K video scenes across 5 spatial scales to create SpaceVista-1M, a dataset comprising approximately 1M spatial QA pairs spanning 19 diverse task types. While specialist models can inject useful domain knowledge, they are not reliable for evaluation. We then build an all-scale benchmark with precise annotations by manually recording, retrieving, and assembling video-based data. However, naive training with SpaceVista-1M often yields suboptimal results due to the potential knowledge conflict. Accordingly, we introduce SpaceVista-7B, a spatial reasoning model that accepts dense inputs beyond semantics and uses scale as an anchor for scale-aware experts and progressive rewards. Finally, extensive evaluations across 5 benchmarks, including our SpaceVista-Bench, demonstrate competitive performance, showcasing strong generalization across all scales and scenarios. Our dataset, model, and benchmark will be released on https://peiwensun2000.github.io/mm2km .
title SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.09606