MiMo-Embodied: X-Embodied Foundation Model Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908996449861632 |
|---|---|
| author | Hao, Xiaoshuai Zhou, Lei Huang, Zhijian Hou, Zhiwen Tang, Yingbo Zhang, Lingfeng Li, Guang Lu, Zheng Ren, Shuhuai Meng, Xianhui Zhang, Yuchen Wu, Jing Lu, Jinghui Dang, Chenxu Guan, Jiayi Wu, Jianhua Hou, Zhiyi Li, Hanbing Xia, Shumeng Zhou, Mingliang Zheng, Yinan Yue, Zihao Gu, Shuhao Tian, Hao Shen, Yuannan Cui, Jianwei Zhang, Wen Xu, Shaoqing Wang, Bing Sun, Haiyang Zhu, Zeyu Jiang, Yuncheng Guo, Zibin Gong, Chuhong Zhang, Chaofan Ding, Wenbo Ma, Kun Chen, Guang Cai, Rui Xiang, Diyun Qu, Heng Luo, Fuli Ye, Hangjun Chen, Long |
| author_facet | Hao, Xiaoshuai Zhou, Lei Huang, Zhijian Hou, Zhiwen Tang, Yingbo Zhang, Lingfeng Li, Guang Lu, Zheng Ren, Shuhuai Meng, Xianhui Zhang, Yuchen Wu, Jing Lu, Jinghui Dang, Chenxu Guan, Jiayi Wu, Jianhua Hou, Zhiyi Li, Hanbing Xia, Shumeng Zhou, Mingliang Zheng, Yinan Yue, Zihao Gu, Shuhao Tian, Hao Shen, Yuannan Cui, Jianwei Zhang, Wen Xu, Shaoqing Wang, Bing Sun, Haiyang Zhu, Zeyu Jiang, Yuncheng Guo, Zibin Gong, Chuhong Zhang, Chaofan Ding, Wenbo Ma, Kun Chen, Guang Cai, Rui Xiang, Diyun Qu, Heng Luo, Fuli Ye, Hangjun Chen, Long |
| contents | We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied AI. MiMo-Embodied sets new records across 17 embodied AI benchmarks in Task Planning, Affordance Prediction and Spatial Understanding, while also excelling in 12 autonomous driving benchmarks across Environmental Perception, Status Prediction, and Driving Planning. Across these tasks, MiMo-Embodied significantly outperforms existing open-source, closed-source, and specialized baselines. Our results indicate that through multi-stage learning, curated data construction, and CoT/RL fine-tuning, these two domains exhibit strong positive transfer and mutually reinforce one another. We provide a detailed analysis of our model design and training methodologies to facilitate further research. Code and models are available at https://github.com/XiaomiMiMo/MiMo-Embodied. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_16518 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MiMo-Embodied: X-Embodied Foundation Model Technical Report Hao, Xiaoshuai Zhou, Lei Huang, Zhijian Hou, Zhiwen Tang, Yingbo Zhang, Lingfeng Li, Guang Lu, Zheng Ren, Shuhuai Meng, Xianhui Zhang, Yuchen Wu, Jing Lu, Jinghui Dang, Chenxu Guan, Jiayi Wu, Jianhua Hou, Zhiyi Li, Hanbing Xia, Shumeng Zhou, Mingliang Zheng, Yinan Yue, Zihao Gu, Shuhao Tian, Hao Shen, Yuannan Cui, Jianwei Zhang, Wen Xu, Shaoqing Wang, Bing Sun, Haiyang Zhu, Zeyu Jiang, Yuncheng Guo, Zibin Gong, Chuhong Zhang, Chaofan Ding, Wenbo Ma, Kun Chen, Guang Cai, Rui Xiang, Diyun Qu, Heng Luo, Fuli Ye, Hangjun Chen, Long Robotics Computation and Language Computer Vision and Pattern Recognition We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied AI. MiMo-Embodied sets new records across 17 embodied AI benchmarks in Task Planning, Affordance Prediction and Spatial Understanding, while also excelling in 12 autonomous driving benchmarks across Environmental Perception, Status Prediction, and Driving Planning. Across these tasks, MiMo-Embodied significantly outperforms existing open-source, closed-source, and specialized baselines. Our results indicate that through multi-stage learning, curated data construction, and CoT/RL fine-tuning, these two domains exhibit strong positive transfer and mutually reinforce one another. We provide a detailed analysis of our model design and training methodologies to facilitate further research. Code and models are available at https://github.com/XiaomiMiMo/MiMo-Embodied. |
| title | MiMo-Embodied: X-Embodied Foundation Model Technical Report |
| topic | Robotics Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.16518 |