Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Meng, Lin, Haokun, Li, Haoyuan, Tang, Haoran, Xu, Rongtao, An, Dong, Liu, Xue, Reid, Ian, Liang, Xiaodan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912753845796864
author Cao, Meng
Lin, Haokun
Li, Haoyuan
Tang, Haoran
Xu, Rongtao
An, Dong
Liu, Xue
Reid, Ian
Liang, Xiaodan
author_facet Cao, Meng
Lin, Haokun
Li, Haoyuan
Tang, Haoran
Xu, Rongtao
An, Dong
Liu, Xue
Reid, Ian
Liang, Xiaodan
contents Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). Current methods predominantly rely on verbal descriptive tuning, which suffers from visual illiteracy, i.e., they learn spatial concepts through textual symbols alone, devoid of connection to their visual manifestations. To bridge this gap, this paper introduces MILO, an Implicit spatIaL wOrld modeling paradigm that simulates human-like spatial imagination. MILO integrates a visual generator to provide geometry-aware feedback, thereby implicitly grounding the MLLM's symbolic reasoning in perceptual experience. Complementing this paradigm, we propose RePE (Relative Positional Encoding), a novel encoding scheme that captures relative camera-pose transformations, offering superior performance over absolute coordinate systems. To support the training, we construct GeoGen, a large-scale Geometry-aware Generative dataset with approximately 2,241 videos and 67,827 observation-action-outcome triplets. Experiments demonstrate that our approach significantly enhances spatial reasoning capabilities across multiple baselines and benchmarks, offering a more holistic understanding of 3D space.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
Cao, Meng
Lin, Haokun
Li, Haoyuan
Tang, Haoran
Xu, Rongtao
An, Dong
Liu, Xue
Reid, Ian
Liang, Xiaodan
Computer Vision and Pattern Recognition
Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). Current methods predominantly rely on verbal descriptive tuning, which suffers from visual illiteracy, i.e., they learn spatial concepts through textual symbols alone, devoid of connection to their visual manifestations. To bridge this gap, this paper introduces MILO, an Implicit spatIaL wOrld modeling paradigm that simulates human-like spatial imagination. MILO integrates a visual generator to provide geometry-aware feedback, thereby implicitly grounding the MLLM's symbolic reasoning in perceptual experience. Complementing this paradigm, we propose RePE (Relative Positional Encoding), a novel encoding scheme that captures relative camera-pose transformations, offering superior performance over absolute coordinate systems. To support the training, we construct GeoGen, a large-scale Geometry-aware Generative dataset with approximately 2,241 videos and 67,827 observation-action-outcome triplets. Experiments demonstrate that our approach significantly enhances spatial reasoning capabilities across multiple baselines and benchmarks, offering a more holistic understanding of 3D space.
title Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01821