RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tai, Cong, Zheng, Zhaoyu, Long, Haixu, Wu, Hansheng, Xiang, Haodong, Long, Zhengbin, Xiong, Jun, Shi, Rong, Zhang, Shizhuang, Qiu, Gang, Wang, He, Li, Ruifeng, Huang, Jun, Chang, Bin, Feng, Shuai, Shen, Tao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908545445789696
author Tai, Cong
Zheng, Zhaoyu
Long, Haixu
Wu, Hansheng
Xiang, Haodong
Long, Zhengbin
Xiong, Jun
Shi, Rong
Zhang, Shizhuang
Qiu, Gang
Wang, He
Li, Ruifeng
Huang, Jun
Chang, Bin
Feng, Shuai
Shen, Tao
author_facet Tai, Cong
Zheng, Zhaoyu
Long, Haixu
Wu, Hansheng
Xiang, Haodong
Long, Zhengbin
Xiong, Jun
Shi, Rong
Zhang, Shizhuang
Qiu, Gang
Wang, He
Li, Ruifeng
Huang, Jun
Chang, Bin
Feng, Shuai
Shen, Tao
contents The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and the real world. To overcome these obstacles, we propose RealMirror, a comprehensive, open-source embodied AI VLA platform. RealMirror builds an efficient, low-cost data collection, model training, and inference system that enables end-to-end VLA research without requiring a real robot. To facilitate model evolution and fair comparison, we also introduce a dedicated VLA benchmark for humanoid robots, featuring multiple scenarios, extensive trajectories, and various VLA models. Furthermore, by integrating generative models and 3D Gaussian Splatting to reconstruct realistic environments and robot models, we successfully demonstrate zero-shot Sim2Real transfer, where models trained exclusively on simulation data can perform tasks on a real robot seamlessly, without any fine-tuning. In conclusion, with the unification of these critical components, RealMirror provides a robust framework that significantly accelerates the development of VLA models for humanoid robots. Project page: https://terminators2025.github.io/RealMirror.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2509_14687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
Tai, Cong
Zheng, Zhaoyu
Long, Haixu
Wu, Hansheng
Xiang, Haodong
Long, Zhengbin
Xiong, Jun
Shi, Rong
Zhang, Shizhuang
Qiu, Gang
Wang, He
Li, Ruifeng
Huang, Jun
Chang, Bin
Feng, Shuai
Shen, Tao
Robotics
The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and the real world. To overcome these obstacles, we propose RealMirror, a comprehensive, open-source embodied AI VLA platform. RealMirror builds an efficient, low-cost data collection, model training, and inference system that enables end-to-end VLA research without requiring a real robot. To facilitate model evolution and fair comparison, we also introduce a dedicated VLA benchmark for humanoid robots, featuring multiple scenarios, extensive trajectories, and various VLA models. Furthermore, by integrating generative models and 3D Gaussian Splatting to reconstruct realistic environments and robot models, we successfully demonstrate zero-shot Sim2Real transfer, where models trained exclusively on simulation data can perform tasks on a real robot seamlessly, without any fine-tuning. In conclusion, with the unification of these critical components, RealMirror provides a robust framework that significantly accelerates the development of VLA models for humanoid robots. Project page: https://terminators2025.github.io/RealMirror.github.io
title RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
topic Robotics
url https://arxiv.org/abs/2509.14687