MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chi, Xiaowei, Ge, Kuangzhi, Liu, Jiaming, Zhou, Siyuan, Jia, Peidong, He, Zichen, Liu, Yuzhen, Li, Tingguang, Han, Lei, Han, Sirui, Zhang, Shanghang, Guo, Yike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916909694320640
author Chi, Xiaowei
Ge, Kuangzhi
Liu, Jiaming
Zhou, Siyuan
Jia, Peidong
He, Zichen
Liu, Yuzhen
Li, Tingguang
Han, Lei
Han, Sirui
Zhang, Shanghang
Guo, Yike
author_facet Chi, Xiaowei
Ge, Kuangzhi
Liu, Jiaming
Zhou, Siyuan
Jia, Peidong
He, Zichen
Liu, Yuzhen
Li, Tingguang
Han, Lei
Han, Sirui
Zhang, Shanghang
Guo, Yike
contents Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling capabilities for predicting future states. Two challenges hinder progress: integrating generative processes into feature learning is both technically and conceptually underdeveloped, and naive frame-by-frame video diffusion is computationally inefficient for real-time robotics. To address these, we propose Manipulate in Dream (MinD), a dual-system world model for real-time, risk-aware planning. MinD uses two asynchronous diffusion processes: a low-frequency visual generator (LoDiff) that predicts future scenes and a high-frequency diffusion policy (HiDiff) that outputs actions. Our key insight is that robotic policies do not require fully denoised frames but can rely on low-resolution latents generated in a single denoising step. To connect early predictions to actions, we introduce DiffMatcher, a video-action alignment module with a novel co-training strategy that synchronizes the two diffusion models. MinD achieves a 63% success rate on RL-Bench, 60% on real-world Franka tasks, and operates at 11.3 FPS, demonstrating the efficiency of single-step latent features for control signals. Furthermore, MinD identifies 74% of potential task failures in advance, providing real-time safety signals for monitoring and intervention. This work establishes a new paradigm for efficient and reliable robotic manipulation using generative world models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
Chi, Xiaowei
Ge, Kuangzhi
Liu, Jiaming
Zhou, Siyuan
Jia, Peidong
He, Zichen
Liu, Yuzhen
Li, Tingguang
Han, Lei
Han, Sirui
Zhang, Shanghang
Guo, Yike
Robotics
Artificial Intelligence
Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling capabilities for predicting future states. Two challenges hinder progress: integrating generative processes into feature learning is both technically and conceptually underdeveloped, and naive frame-by-frame video diffusion is computationally inefficient for real-time robotics. To address these, we propose Manipulate in Dream (MinD), a dual-system world model for real-time, risk-aware planning. MinD uses two asynchronous diffusion processes: a low-frequency visual generator (LoDiff) that predicts future scenes and a high-frequency diffusion policy (HiDiff) that outputs actions. Our key insight is that robotic policies do not require fully denoised frames but can rely on low-resolution latents generated in a single denoising step. To connect early predictions to actions, we introduce DiffMatcher, a video-action alignment module with a novel co-training strategy that synchronizes the two diffusion models. MinD achieves a 63% success rate on RL-Bench, 60% on real-world Franka tasks, and operates at 11.3 FPS, demonstrating the efficiency of single-step latent features for control signals. Furthermore, MinD identifies 74% of potential task failures in advance, providing real-time safety signals for monitoring and intervention. This work establishes a new paradigm for efficient and reliable robotic manipulation using generative world models.
title MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2506.18897