Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Pingyue, Huang, Zihan, Wang, Yue, Zhang, Jieyu, Xue, Letian, Wang, Zihan, Wang, Qineng, Chandrasegaran, Keshigeyan, Zhang, Ruohan, Choi, Yejin, Krishna, Ranjay, Wu, Jiajun, Fei-Fei, Li, Li, Manling
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908818328256512
author Zhang, Pingyue
Huang, Zihan
Wang, Yue
Zhang, Jieyu
Xue, Letian
Wang, Zihan
Wang, Qineng
Chandrasegaran, Keshigeyan
Zhang, Ruohan
Choi, Yejin
Krishna, Ranjay
Wu, Jiajun
Fei-Fei, Li
Li, Manling
author_facet Zhang, Pingyue
Huang, Zihan
Wang, Yue
Zhang, Jieyu
Xue, Letian
Wang, Zihan
Wang, Qineng
Chandrasegaran, Keshigeyan
Zhang, Ruohan
Choi, Yejin
Krishna, Ranjay
Wu, Jiajun
Fei-Fei, Li
Li, Manling
contents Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploration remains understudied. We propose Theory of Space, defined as an agent's ability to actively acquire information through self-directed, active exploration and to construct, revise, and exploit a spatial belief from sequential, partial observations. We evaluate this through a benchmark where the goal is curiosity-driven exploration to build an accurate cognitive map. A key innovation is spatial belief probing, which prompts models to reveal their internal spatial representations at each step. Our evaluation of state-of-the-art models reveals several critical bottlenecks. First, we identify an Active-Passive Gap, where performance drops significantly when agents must autonomously gather information. Second, we find high inefficiency, as models explore unsystematically compared to program-based proxies. Through belief probing, we diagnose that while perception is an initial bottleneck, global beliefs suffer from instability that causes spatial knowledge to degrade over time. Finally, using a false belief paradigm, we uncover Belief Inertia, where agents fail to update obsolete priors with new evidence. This issue is present in text-based agents but is particularly severe in vision-based models. Our findings suggest that current foundation models struggle to maintain coherent, revisable spatial beliefs during active exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07055
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
Zhang, Pingyue
Huang, Zihan
Wang, Yue
Zhang, Jieyu
Xue, Letian
Wang, Zihan
Wang, Qineng
Chandrasegaran, Keshigeyan
Zhang, Ruohan
Choi, Yejin
Krishna, Ranjay
Wu, Jiajun
Fei-Fei, Li
Li, Manling
Artificial Intelligence
Computation and Language
Machine Learning
Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploration remains understudied. We propose Theory of Space, defined as an agent's ability to actively acquire information through self-directed, active exploration and to construct, revise, and exploit a spatial belief from sequential, partial observations. We evaluate this through a benchmark where the goal is curiosity-driven exploration to build an accurate cognitive map. A key innovation is spatial belief probing, which prompts models to reveal their internal spatial representations at each step. Our evaluation of state-of-the-art models reveals several critical bottlenecks. First, we identify an Active-Passive Gap, where performance drops significantly when agents must autonomously gather information. Second, we find high inefficiency, as models explore unsystematically compared to program-based proxies. Through belief probing, we diagnose that while perception is an initial bottleneck, global beliefs suffer from instability that causes spatial knowledge to degrade over time. Finally, using a false belief paradigm, we uncover Belief Inertia, where agents fail to update obsolete priors with new evidence. This issue is present in text-based agents but is particularly severe in vision-based models. Our findings suggest that current foundation models struggle to maintain coherent, revisable spatial beliefs during active exploration.
title Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.07055