MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916909694320640 |
|---|---|
| author | Chi, Xiaowei Ge, Kuangzhi Liu, Jiaming Zhou, Siyuan Jia, Peidong He, Zichen Liu, Yuzhen Li, Tingguang Han, Lei Han, Sirui Zhang, Shanghang Guo, Yike |
| author_facet | Chi, Xiaowei Ge, Kuangzhi Liu, Jiaming Zhou, Siyuan Jia, Peidong He, Zichen Liu, Yuzhen Li, Tingguang Han, Lei Han, Sirui Zhang, Shanghang Guo, Yike |
| contents | Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling capabilities for predicting future states. Two challenges hinder progress: integrating generative processes into feature learning is both technically and conceptually underdeveloped, and naive frame-by-frame video diffusion is computationally inefficient for real-time robotics. To address these, we propose Manipulate in Dream (MinD), a dual-system world model for real-time, risk-aware planning. MinD uses two asynchronous diffusion processes: a low-frequency visual generator (LoDiff) that predicts future scenes and a high-frequency diffusion policy (HiDiff) that outputs actions. Our key insight is that robotic policies do not require fully denoised frames but can rely on low-resolution latents generated in a single denoising step. To connect early predictions to actions, we introduce DiffMatcher, a video-action alignment module with a novel co-training strategy that synchronizes the two diffusion models. MinD achieves a 63% success rate on RL-Bench, 60% on real-world Franka tasks, and operates at 11.3 FPS, demonstrating the efficiency of single-step latent features for control signals. Furthermore, MinD identifies 74% of potential task failures in advance, providing real-time safety signals for monitoring and intervention. This work establishes a new paradigm for efficient and reliable robotic manipulation using generative world models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18897 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis Chi, Xiaowei Ge, Kuangzhi Liu, Jiaming Zhou, Siyuan Jia, Peidong He, Zichen Liu, Yuzhen Li, Tingguang Han, Lei Han, Sirui Zhang, Shanghang Guo, Yike Robotics Artificial Intelligence Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling capabilities for predicting future states. Two challenges hinder progress: integrating generative processes into feature learning is both technically and conceptually underdeveloped, and naive frame-by-frame video diffusion is computationally inefficient for real-time robotics. To address these, we propose Manipulate in Dream (MinD), a dual-system world model for real-time, risk-aware planning. MinD uses two asynchronous diffusion processes: a low-frequency visual generator (LoDiff) that predicts future scenes and a high-frequency diffusion policy (HiDiff) that outputs actions. Our key insight is that robotic policies do not require fully denoised frames but can rely on low-resolution latents generated in a single denoising step. To connect early predictions to actions, we introduce DiffMatcher, a video-action alignment module with a novel co-training strategy that synchronizes the two diffusion models. MinD achieves a 63% success rate on RL-Bench, 60% on real-world Franka tasks, and operates at 11.3 FPS, demonstrating the efficiency of single-step latent features for control signals. Furthermore, MinD identifies 74% of potential task failures in advance, providing real-time safety signals for monitoring and intervention. This work establishes a new paradigm for efficient and reliable robotic manipulation using generative world models. |
| title | MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis |
| topic | Robotics Artificial Intelligence |
| url | https://arxiv.org/abs/2506.18897 |