Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Yibin, Xu, Wang, Zhang, Wanyue, Zhi, Helu, Huang, Jingjing, Xu, Yangbin, Sun, Yangang, Zhu, Conghui, Zhao, Tiejun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917320456142848
author Huang, Yibin
Xu, Wang
Zhang, Wanyue
Zhi, Helu
Huang, Jingjing
Xu, Yangbin
Sun, Yangang
Zhu, Conghui
Zhao, Tiejun
author_facet Huang, Yibin
Xu, Wang
Zhang, Wanyue
Zhi, Helu
Huang, Jingjing
Xu, Yangbin
Sun, Yangang
Zhu, Conghui
Zhao, Tiejun
contents Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial understanding via grid-based cognitive maps. However, current grid-based map methods rely on discretized representations, which limit the model's ability in fine-grained spatial reasoning. To overcome this limitation, we propose Video2Layout, a framework for reconstructing metric-grounded spatial layouts from video. The framework uses continuous object boundary coordinates to enable quantitative spatial computation, which effectively reduces ambiguity in natural language descriptions of spatial relationships. Specifically, our method comprises two stages. First, in supervised fine-tuning stage, we construct a high-quality dataset from the AI2THOR simulator, which enables the model to learn the mapping from visual inputs to precise boundary coordinates. Subsequently, a reinforcement fine-tuning stage enhances the model's real-world generalization capabilities. Based on the above framework, we investigate factors that affect cognitive map accuracy and quantify its relationship with task performance. Evaluated on mainstream spatial reasoning benchmarks, our model, V2LO-7B, achieves an average improvement of 3.24\% over the model trained on grid maps, validating the superiority of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
Huang, Yibin
Xu, Wang
Zhang, Wanyue
Zhi, Helu
Huang, Jingjing
Xu, Yangbin
Sun, Yangang
Zhu, Conghui
Zhao, Tiejun
Computer Vision and Pattern Recognition
Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial understanding via grid-based cognitive maps. However, current grid-based map methods rely on discretized representations, which limit the model's ability in fine-grained spatial reasoning. To overcome this limitation, we propose Video2Layout, a framework for reconstructing metric-grounded spatial layouts from video. The framework uses continuous object boundary coordinates to enable quantitative spatial computation, which effectively reduces ambiguity in natural language descriptions of spatial relationships. Specifically, our method comprises two stages. First, in supervised fine-tuning stage, we construct a high-quality dataset from the AI2THOR simulator, which enables the model to learn the mapping from visual inputs to precise boundary coordinates. Subsequently, a reinforcement fine-tuning stage enhances the model's real-world generalization capabilities. Based on the above framework, we investigate factors that affect cognitive map accuracy and quantify its relationship with task performance. Evaluated on mainstream spatial reasoning benchmarks, our model, V2LO-7B, achieves an average improvement of 3.24\% over the model trained on grid maps, validating the superiority of our method.
title Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16160