RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yu, Chao, Wang, Yuanqing, Guo, Zhen, Lin, Hao, Xu, Si, Zang, Hongzhi, Zhang, Quanlu, Wu, Yongji, Zhu, Chunyang, Hu, Junhao, Huang, Zixiao, Wei, Mingjie, Xie, Yuqing, Yang, Ke, Dai, Bo, Xu, Zhexuan, Du, Jiakun, Wang, Xiangyuan, Fu, Xu, Shi, Letong, Liu, Zhihao, Chen, Kang, Liu, Weilin, Liu, Gang, Li, Boxun, Yang, Jianlei, Yang, Zhi, Dai, Guohao, Wang, Yu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918264891768832
author Yu, Chao
Wang, Yuanqing
Guo, Zhen
Lin, Hao
Xu, Si
Zang, Hongzhi
Zhang, Quanlu
Wu, Yongji
Zhu, Chunyang
Hu, Junhao
Huang, Zixiao
Wei, Mingjie
Xie, Yuqing
Yang, Ke
Dai, Bo
Xu, Zhexuan
Du, Jiakun
Wang, Xiangyuan
Fu, Xu
Shi, Letong
Liu, Zhihao
Chen, Kang
Liu, Weilin
Liu, Gang
Li, Boxun
Yang, Jianlei
Yang, Zhi
Dai, Guohao
Wang, Yu
author_facet Yu, Chao
Wang, Yuanqing
Guo, Zhen
Lin, Hao
Xu, Si
Zang, Hongzhi
Zhang, Quanlu
Wu, Yongji
Zhu, Chunyang
Hu, Junhao
Huang, Zixiao
Wei, Mingjie
Xie, Yuqing
Yang, Ke
Dai, Bo
Xu, Zhexuan
Du, Jiakun
Wang, Xiangyuan
Fu, Xu
Shi, Letong
Liu, Zhihao
Chen, Kang
Liu, Weilin
Liu, Gang
Li, Boxun
Yang, Jianlei
Yang, Zhi
Dai, Guohao
Wang, Yu
contents Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\times-2.43\times$ speedup in end-to-end training throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15965
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
Yu, Chao
Wang, Yuanqing
Guo, Zhen
Lin, Hao
Xu, Si
Zang, Hongzhi
Zhang, Quanlu
Wu, Yongji
Zhu, Chunyang
Hu, Junhao
Huang, Zixiao
Wei, Mingjie
Xie, Yuqing
Yang, Ke
Dai, Bo
Xu, Zhexuan
Du, Jiakun
Wang, Xiangyuan
Fu, Xu
Shi, Letong
Liu, Zhihao
Chen, Kang
Liu, Weilin
Liu, Gang
Li, Boxun
Yang, Jianlei
Yang, Zhi
Dai, Guohao
Wang, Yu
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\times-2.43\times$ speedup in end-to-end training throughput.
title RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.15965