RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914355170246656 |
|---|---|
| author | Hou, Chengkai Wu, Kun Liu, Jiaming Che, Zhengping Wu, Di Liao, Fei Li, Guangrun He, Jingyang Feng, Qiuxuan Jin, Zhao Gu, Chenyang Liu, Zhuoyang Han, Nuowei Mi, Xiangju Lv, Yaoxu Fu, Yankai Dai, Gaole Gu, Langzhe Li, Tao Zhang, Yuheng Zhang, Yixue Wang, Xinhua Fan, Shichao Li, Meng Zhao, Zhen Liu, Ning Xu, Zhiyuan Ren, Pei Ji, Junjie Liu, Haonan Cheng, Kuan Zhang, Shanghang Tang, Jian |
| author_facet | Hou, Chengkai Wu, Kun Liu, Jiaming Che, Zhengping Wu, Di Liao, Fei Li, Guangrun He, Jingyang Feng, Qiuxuan Jin, Zhao Gu, Chenyang Liu, Zhuoyang Han, Nuowei Mi, Xiangju Lv, Yaoxu Fu, Yankai Dai, Gaole Gu, Langzhe Li, Tao Zhang, Yuheng Zhang, Yixue Wang, Xinhua Fan, Shichao Li, Meng Zhao, Zhen Liu, Ning Xu, Zhiyuan Ren, Pei Ji, Junjie Liu, Haonan Cheng, Kuan Zhang, Shanghang Tang, Jian |
| contents | While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to generalize across long-horizon bimanual tasks and mobile manipulation in unstructured environments remains limited. To bridge this gap, we present RoboMIND 2.0, a comprehensive real-world dataset comprising over 310K dual-arm manipulation trajectories collected across six distinct robot embodiments and 739 complex tasks. Crucially, to support research in contact-rich and spatially extended tasks, the dataset incorporates 12K tactile-enhanced episodes and 20K mobile manipulation trajectories. Complementing this physical data, we construct high-fidelity digital twins of our real-world environments, releasing an additional 20K-trajectory simulated dataset to facilitate robust sim-to-real transfer. To fully exploit the potential of RoboMIND 2.0, we propose MIND-2 system, a hierarchical dual-system frame-work optimized via offline reinforcement learning. MIND-2 integrates a high-level semantic planner (MIND-2-VLM) to decompose abstract natural language instructions into grounded subgoals, coupled with a low-level Vision-Language-Action executor (MIND-2-VLA), which generates precise, proprioception-aware motor actions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_24653 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence Hou, Chengkai Wu, Kun Liu, Jiaming Che, Zhengping Wu, Di Liao, Fei Li, Guangrun He, Jingyang Feng, Qiuxuan Jin, Zhao Gu, Chenyang Liu, Zhuoyang Han, Nuowei Mi, Xiangju Lv, Yaoxu Fu, Yankai Dai, Gaole Gu, Langzhe Li, Tao Zhang, Yuheng Zhang, Yixue Wang, Xinhua Fan, Shichao Li, Meng Zhao, Zhen Liu, Ning Xu, Zhiyuan Ren, Pei Ji, Junjie Liu, Haonan Cheng, Kuan Zhang, Shanghang Tang, Jian Robotics While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to generalize across long-horizon bimanual tasks and mobile manipulation in unstructured environments remains limited. To bridge this gap, we present RoboMIND 2.0, a comprehensive real-world dataset comprising over 310K dual-arm manipulation trajectories collected across six distinct robot embodiments and 739 complex tasks. Crucially, to support research in contact-rich and spatially extended tasks, the dataset incorporates 12K tactile-enhanced episodes and 20K mobile manipulation trajectories. Complementing this physical data, we construct high-fidelity digital twins of our real-world environments, releasing an additional 20K-trajectory simulated dataset to facilitate robust sim-to-real transfer. To fully exploit the potential of RoboMIND 2.0, we propose MIND-2 system, a hierarchical dual-system frame-work optimized via offline reinforcement learning. MIND-2 integrates a high-level semantic planner (MIND-2-VLM) to decompose abstract natural language instructions into grounded subgoals, coupled with a low-level Vision-Language-Action executor (MIND-2-VLA), which generates precise, proprioception-aware motor actions. |
| title | RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence |
| topic | Robotics |
| url | https://arxiv.org/abs/2512.24653 |