HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yongjun, Zhang, Shuai, Gai, Jiading, Zhang, Xiyuan, Han, Boran, Wang, Bernie, Rangwala, Huzefa, Karypis, George
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911583495520256
author He, Yongjun
Zhang, Shuai
Gai, Jiading
Zhang, Xiyuan
Han, Boran
Wang, Bernie
Rangwala, Huzefa
Karypis, George
author_facet He, Yongjun
Zhang, Shuai
Gai, Jiading
Zhang, Xiyuan
Han, Boran
Wang, Bernie
Rangwala, Huzefa
Karypis, George
contents As large language models (LLMs) continue to scale and new GPUs are released even more frequently, there is an increasing demand for LLM post-training in heterogeneous environments to fully leverage underutilized mid-range or previous-generation GPUs and alleviate the shortage of homogeneous high-end GPUs within a single availability zone. However, achieving high-performance reinforcement learning (RL) training for LLMs on such computing resources remains challenging because the workflow involves multiple models and tasks with complex computation and data dependencies. In this paper, we present HetRL, a distributed system for efficient RL training in infrastructures with heterogeneous GPUs and networks. HetRL formulates the scheduling of RL training in heterogeneous environments as a constrained joint optimization problem and provides two complementary approaches for addressing this problem: (1) a hybrid scheduling algorithm that efficiently identifies near-optimal solutions, and (2) an integer linear programming (ILP)-based scheduling algorithm that obtains optimal solutions, enabling flexible trade-offs between solution optimality and efficiency. Our extensive evaluation, consuming 20,000 GPU-hours, shows that HetRL achieves up to 9.17x the throughput of state-of-the-art systems, and 3.17x on average, across a wide range of workloads and settings.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
He, Yongjun
Zhang, Shuai
Gai, Jiading
Zhang, Xiyuan
Han, Boran
Wang, Bernie
Rangwala, Huzefa
Karypis, George
Distributed, Parallel, and Cluster Computing
As large language models (LLMs) continue to scale and new GPUs are released even more frequently, there is an increasing demand for LLM post-training in heterogeneous environments to fully leverage underutilized mid-range or previous-generation GPUs and alleviate the shortage of homogeneous high-end GPUs within a single availability zone. However, achieving high-performance reinforcement learning (RL) training for LLMs on such computing resources remains challenging because the workflow involves multiple models and tasks with complex computation and data dependencies. In this paper, we present HetRL, a distributed system for efficient RL training in infrastructures with heterogeneous GPUs and networks. HetRL formulates the scheduling of RL training in heterogeneous environments as a constrained joint optimization problem and provides two complementary approaches for addressing this problem: (1) a hybrid scheduling algorithm that efficiently identifies near-optimal solutions, and (2) an integer linear programming (ILP)-based scheduling algorithm that obtains optimal solutions, enabling flexible trade-offs between solution optimality and efficiency. Our extensive evaluation, consuming 20,000 GPU-hours, shows that HetRL achieves up to 9.17x the throughput of state-of-the-art systems, and 3.17x on average, across a wide range of workloads and settings.
title HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.12476