MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yuncong, Liu, Jiageng, Zhang, Zheyuan, Zhou, Siyuan, Tan, Reuben, Yang, Jianwei, Du, Yilun, Gan, Chuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning 3D Persistent Embodied World Models
by: Zhou, Siyuan, et al.
Published: (2025)
by: Zhou, Siyuan, et al.
Published: (2025)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024)
by: Yang, Yuncong, et al.
Published: (2024)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
TesserAct: Learning 4D Embodied World Models
by: Zhen, Haoyu, et al.
Published: (2025)
by: Zhen, Haoyu, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Fast Spatial Memory with Elastic Test-Time Training
by: Ma, Ziqiao, et al.
Published: (2026)
by: Ma, Ziqiao, et al.
Published: (2026)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
by: Qi, Han, et al.
Published: (2025)
by: Qi, Han, et al.
Published: (2025)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
Virtual Community: An Open World for Humans, Robots, and Society
by: Zhou, Qinhong, et al.
Published: (2025)
by: Zhou, Qinhong, et al.
Published: (2025)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Building Cooperative Embodied Agents Modularly with Large Language Models
by: Zhang, Hongxin, et al.
Published: (2023)
by: Zhang, Hongxin, et al.
Published: (2023)
RoboDreamer: Learning Compositional World Models for Robot Imagination
by: Zhou, Siyuan, et al.
Published: (2024)
by: Zhou, Siyuan, et al.
Published: (2024)
SimScale: Learning to Drive via Real-World Simulation at Scale
by: Tian, Haochen, et al.
Published: (2025)
by: Tian, Haochen, et al.
Published: (2025)
3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
by: Zhi, Hongyan, et al.
Published: (2025)
by: Zhi, Hongyan, et al.
Published: (2025)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
by: Lu, Yiren, et al.
Published: (2026)
by: Lu, Yiren, et al.
Published: (2026)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
Robot Learning from a Physical World Model
by: Mao, Jiageng, et al.
Published: (2025)
by: Mao, Jiageng, et al.
Published: (2025)
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
by: Che, Lirong, et al.
Published: (2026)
by: Che, Lirong, et al.
Published: (2026)
World Model for Robot Learning: A Comprehensive Survey
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2026)
by: Shang, Shuyao, et al.
Published: (2026)
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
by: Qin, Yiran, et al.
Published: (2026)
by: Qin, Yiran, et al.
Published: (2026)
Compositional Generative Modeling: A Single Model is Not All You Need
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
by: Meng, Siyuan, et al.
Published: (2026)
by: Meng, Siyuan, et al.
Published: (2026)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
CoNav: A Benchmark for Human-Centered Collaborative Navigation
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
by: Liao, Yue, et al.
Published: (2025)
by: Liao, Yue, et al.
Published: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
by: Li, Sibo, et al.
Published: (2025)
by: Li, Sibo, et al.
Published: (2025)
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
by: Jing, Bowen, et al.
Published: (2026)
by: Jing, Bowen, et al.
Published: (2026)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
by: Yang, Dejie, et al.
Published: (2025)
by: Yang, Dejie, et al.
Published: (2025)
Can Transformers Capture Spatial Relations between Objects?
by: Wen, Chuan, et al.
Published: (2024)
by: Wen, Chuan, et al.
Published: (2024)
CoPeD-Advancing Multi-Robot Collaborative Perception: A Comprehensive Dataset in Real-World Environments
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
Similar Items
-
Learning 3D Persistent Embodied World Models
by: Zhou, Siyuan, et al.
Published: (2025) -
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024) -
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025) -
TesserAct: Learning 4D Embodied World Models
by: Zhen, Haoyu, et al.
Published: (2025) -
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)