MindJourney: Test-Time Scaling with World Models for Spatial Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yuncong, Liu, Jiageng, Zhang, Zheyuan, Zhou, Siyuan, Tan, Reuben, Yang, Jianwei, Du, Yilun, Gan, Chuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912682461888512
author Yang, Yuncong
Liu, Jiageng
Zhang, Zheyuan
Zhou, Siyuan
Tan, Reuben
Yang, Jianwei
Du, Yilun
Gan, Chuang
author_facet Yang, Yuncong
Liu, Jiageng
Zhang, Zheyuan
Zhou, Siyuan
Tan, Reuben
Yang, Jianwei
Du, Yilun
Gan, Chuang
contents Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision-language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: they perceive 2D images but lack an internal model of 3D dynamics. We therefore propose MindJourney, a test-time scaling framework that grants a VLM with this missing capability by coupling it to a controllable world model based on video diffusion. The VLM iteratively sketches a concise camera trajectory, while the world model synthesizes the corresponding view at each step. The VLM then reasons over this multi-view evidence gathered during the interactive exploration. Without any fine-tuning, our MindJourney achieves over an average 7.7% performance boost on the representative spatial reasoning benchmark SAT, showing that pairing VLMs with world models for test-time scaling offers a simple, plug-and-play route to robust 3D reasoning. Meanwhile, our method also improves upon the test-time inference VLMs trained through reinforcement learning, which demonstrates the potential of our method that utilizes world models for test-time scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
Yang, Yuncong
Liu, Jiageng
Zhang, Zheyuan
Zhou, Siyuan
Tan, Reuben
Yang, Jianwei
Du, Yilun
Gan, Chuang
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision-language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: they perceive 2D images but lack an internal model of 3D dynamics. We therefore propose MindJourney, a test-time scaling framework that grants a VLM with this missing capability by coupling it to a controllable world model based on video diffusion. The VLM iteratively sketches a concise camera trajectory, while the world model synthesizes the corresponding view at each step. The VLM then reasons over this multi-view evidence gathered during the interactive exploration. Without any fine-tuning, our MindJourney achieves over an average 7.7% performance boost on the representative spatial reasoning benchmark SAT, showing that pairing VLMs with world models for test-time scaling offers a simple, plug-and-play route to robust 3D reasoning. Meanwhile, our method also improves upon the test-time inference VLMs trained through reinforcement learning, which demonstrates the potential of our method that utilizes world models for test-time scaling.
title MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2507.12508