_version_ 1866916063764021248
author Jia, Yufei
Cao, Zhanxiang
Yu, Mingrui
Zhang, Heng
Chen, Shenyu
Jiang, Dixuan
Li, Meng
Li, Xiaofan
Liu, Yiyang
Wu, Junzhe
Li, Zheng
Fang, XiLin
Cui, Tingyu
Fu, Shengcheng
Li, Haoyang
Wang, Anqi
Wang, Zifan
Zhu, Dongjie
Cao, Chenyu
Huang, Zhenbiao
Zheng, Ziang
Lu, Jie
Ma, Xin
Wei, Zhengyang
Zhao, Xiang
Zhan, Tianyue
He, Ye
Chen, Yuxiang
Jiang, Yizhou
Li, Yue
Ge, Haizhou
Dong, Yuhang
Jia, Fan
Zhang, Ziheng
Zhang, Meng
Deng, Xiwa
Chen, Zhixing
Shao, Hanyang
Dong, Chenxin
Li, Yixuan
Chen, Yizhi
Chen, Bokui
Zhang, Kaifeng
Cui, Hanqing
Qin, Yusen
Huang, Ruqi
Han, Lei
Wang, Tiancai
Li, Xiang
Gao, Yue
Zhou, Guyue
author_facet Jia, Yufei
Cao, Zhanxiang
Yu, Mingrui
Zhang, Heng
Chen, Shenyu
Jiang, Dixuan
Li, Meng
Li, Xiaofan
Liu, Yiyang
Wu, Junzhe
Li, Zheng
Fang, XiLin
Cui, Tingyu
Fu, Shengcheng
Li, Haoyang
Wang, Anqi
Wang, Zifan
Zhu, Dongjie
Cao, Chenyu
Huang, Zhenbiao
Zheng, Ziang
Lu, Jie
Ma, Xin
Wei, Zhengyang
Zhao, Xiang
Zhan, Tianyue
He, Ye
Chen, Yuxiang
Jiang, Yizhou
Li, Yue
Ge, Haizhou
Dong, Yuhang
Jia, Fan
Zhang, Ziheng
Zhang, Meng
Deng, Xiwa
Chen, Zhixing
Shao, Hanyang
Dong, Chenxin
Li, Yixuan
Chen, Yizhi
Chen, Bokui
Zhang, Kaifeng
Cui, Hanqing
Qin, Yusen
Huang, Ruqi
Han, Lei
Wang, Tiancai
Li, Xiang
Gao, Yue
Zhou, Guyue
contents Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-centric execution path. This paradigm has greatly improved training speed, but it has also encouraged a default assumption that efficient training requires physics to reside on the GPU. We revisit this assumption. Our view is that, in simulation-dominated robot control, the essential question is not which processor runs physics, but whether simulation throughput, policy learning, and runtime synchronization form an efficient end-to-end loop. We present UniLab, a heterogeneous CPU-simulation / GPU-learning architecture that decouples CPU-parallel simulation from GPU policy updates through a unified runtime for data movement, buffering, and synchronization. UniLab is implemented as a complete and extensible training system using MuJoCoUni and MotrixSim CPU-batched physics backends, supporting PPO, FastSAC, FlashSAC, and APPO. On representative simulation-based robot control tasks, UniLab improves end-to-end training efficiency by 3--10$\times$ under the same hardware configuration, while reducing dependence on the NVIDIA CUDA-based software stack and supporting cross-platform execution on the Apple macOS platform and the AMD ROCm and Intel XPU accelerator backends. These results show that GPU simulation is an effective path to efficient training, but not a necessary one, broadening the practical system choices available for robot RL training. Project page: https://unilabsim.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30313
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
Jia, Yufei
Cao, Zhanxiang
Yu, Mingrui
Zhang, Heng
Chen, Shenyu
Jiang, Dixuan
Li, Meng
Li, Xiaofan
Liu, Yiyang
Wu, Junzhe
Li, Zheng
Fang, XiLin
Cui, Tingyu
Fu, Shengcheng
Li, Haoyang
Wang, Anqi
Wang, Zifan
Zhu, Dongjie
Cao, Chenyu
Huang, Zhenbiao
Zheng, Ziang
Lu, Jie
Ma, Xin
Wei, Zhengyang
Zhao, Xiang
Zhan, Tianyue
He, Ye
Chen, Yuxiang
Jiang, Yizhou
Li, Yue
Ge, Haizhou
Dong, Yuhang
Jia, Fan
Zhang, Ziheng
Zhang, Meng
Deng, Xiwa
Chen, Zhixing
Shao, Hanyang
Dong, Chenxin
Li, Yixuan
Chen, Yizhi
Chen, Bokui
Zhang, Kaifeng
Cui, Hanqing
Qin, Yusen
Huang, Ruqi
Han, Lei
Wang, Tiancai
Li, Xiang
Gao, Yue
Zhou, Guyue
Robotics
68T40
I.2.9
Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-centric execution path. This paradigm has greatly improved training speed, but it has also encouraged a default assumption that efficient training requires physics to reside on the GPU. We revisit this assumption. Our view is that, in simulation-dominated robot control, the essential question is not which processor runs physics, but whether simulation throughput, policy learning, and runtime synchronization form an efficient end-to-end loop. We present UniLab, a heterogeneous CPU-simulation / GPU-learning architecture that decouples CPU-parallel simulation from GPU policy updates through a unified runtime for data movement, buffering, and synchronization. UniLab is implemented as a complete and extensible training system using MuJoCoUni and MotrixSim CPU-batched physics backends, supporting PPO, FastSAC, FlashSAC, and APPO. On representative simulation-based robot control tasks, UniLab improves end-to-end training efficiency by 3--10$\times$ under the same hardware configuration, while reducing dependence on the NVIDIA CUDA-based software stack and supporting cross-platform execution on the Apple macOS platform and the AMD ROCm and Intel XPU accelerator backends. These results show that GPU simulation is an effective path to efficient training, but not a necessary one, broadening the practical system choices available for robot RL training. Project page: https://unilabsim.github.io.
title UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
topic Robotics
68T40
I.2.9
url https://arxiv.org/abs/2605.30313