MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruicheng, Zhang, Mingyang, Zhou, Jun, Guo, Zhangrui, Xu, Zunnan, Liu, Xiaofan, Zhong, Zhizhou, Yan, Puxin, Luo, Haocheng, Li, Xiu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912964141907968
author Zhang, Ruicheng
Zhang, Mingyang
Zhou, Jun
Guo, Zhangrui
Xu, Zunnan
Liu, Xiaofan
Zhong, Zhizhou
Yan, Puxin
Luo, Haocheng
Li, Xiu
author_facet Zhang, Ruicheng
Zhang, Mingyang
Zhou, Jun
Guo, Zhangrui
Xu, Zunnan
Liu, Xiaofan
Zhong, Zhizhou
Yan, Puxin
Luo, Haocheng
Li, Xiu
contents Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models in this domain are limited to synthesizing short clips of simple actions and often rely on manually defined trajectories. To this end, we introduce MIND-V, a cognitive hierarchical world model designed to synthesize physically plausible and logically coherent videos of long-horizon robotic manipulation. Inspired by cognitive science, MIND-V bridges high-level reasoning with pixel-level synthesis through three core components: a Semantic Reasoning Hub (SRH) that leverages a pre-trained vision-language model for task planning; a Behavioral Semantic Bridge (BSB) that translates abstract instructions into domain-invariant representations; and a Motor Video Generator (MVG) for conditional video rendering. MIND-V employs Staged Visual Future Rollouts, a test-time optimization strategy to enhance long-horizon robustness. To enforce adherence to physical laws, we introduce a GRPO reinforcement learning post-training phase guided by a novel Physical Foresight Coherence (PFC) reward. PFC leverages the V-JEPA2 world model as a physics referee to penalize implausible dynamics in the latent feature space. Experiments confirm MIND-V's SOTA performance in long-horizon simulation and its significant value for policy learning, introducing a scalable and fully autonomous framework for embodied data synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06628
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
Zhang, Ruicheng
Zhang, Mingyang
Zhou, Jun
Guo, Zhangrui
Xu, Zunnan
Liu, Xiaofan
Zhong, Zhizhou
Yan, Puxin
Luo, Haocheng
Li, Xiu
Robotics
Computer Vision and Pattern Recognition
Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models in this domain are limited to synthesizing short clips of simple actions and often rely on manually defined trajectories. To this end, we introduce MIND-V, a cognitive hierarchical world model designed to synthesize physically plausible and logically coherent videos of long-horizon robotic manipulation. Inspired by cognitive science, MIND-V bridges high-level reasoning with pixel-level synthesis through three core components: a Semantic Reasoning Hub (SRH) that leverages a pre-trained vision-language model for task planning; a Behavioral Semantic Bridge (BSB) that translates abstract instructions into domain-invariant representations; and a Motor Video Generator (MVG) for conditional video rendering. MIND-V employs Staged Visual Future Rollouts, a test-time optimization strategy to enhance long-horizon robustness. To enforce adherence to physical laws, we introduce a GRPO reinforcement learning post-training phase guided by a novel Physical Foresight Coherence (PFC) reward. PFC leverages the V-JEPA2 world model as a physics referee to penalize implausible dynamics in the latent feature space. Experiments confirm MIND-V's SOTA performance in long-horizon simulation and its significant value for policy learning, introducing a scalable and fully autonomous framework for embodied data synthesis.
title MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06628