RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918264891768832 |
|---|---|
| author | Yu, Chao Wang, Yuanqing Guo, Zhen Lin, Hao Xu, Si Zang, Hongzhi Zhang, Quanlu Wu, Yongji Zhu, Chunyang Hu, Junhao Huang, Zixiao Wei, Mingjie Xie, Yuqing Yang, Ke Dai, Bo Xu, Zhexuan Du, Jiakun Wang, Xiangyuan Fu, Xu Shi, Letong Liu, Zhihao Chen, Kang Liu, Weilin Liu, Gang Li, Boxun Yang, Jianlei Yang, Zhi Dai, Guohao Wang, Yu |
| author_facet | Yu, Chao Wang, Yuanqing Guo, Zhen Lin, Hao Xu, Si Zang, Hongzhi Zhang, Quanlu Wu, Yongji Zhu, Chunyang Hu, Junhao Huang, Zixiao Wei, Mingjie Xie, Yuqing Yang, Ke Dai, Bo Xu, Zhexuan Du, Jiakun Wang, Xiangyuan Fu, Xu Shi, Letong Liu, Zhihao Chen, Kang Liu, Weilin Liu, Gang Li, Boxun Yang, Jianlei Yang, Zhi Dai, Guohao Wang, Yu |
| contents | Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\times-2.43\times$ speedup in end-to-end training throughput. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15965 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation Yu, Chao Wang, Yuanqing Guo, Zhen Lin, Hao Xu, Si Zang, Hongzhi Zhang, Quanlu Wu, Yongji Zhu, Chunyang Hu, Junhao Huang, Zixiao Wei, Mingjie Xie, Yuqing Yang, Ke Dai, Bo Xu, Zhexuan Du, Jiakun Wang, Xiangyuan Fu, Xu Shi, Letong Liu, Zhihao Chen, Kang Liu, Weilin Liu, Gang Li, Boxun Yang, Jianlei Yang, Zhi Dai, Guohao Wang, Yu Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\times-2.43\times$ speedup in end-to-end training throughput. |
| title | RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation |
| topic | Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2509.15965 |