RL in the Wild: Characterizing RLVR Training in LLM Deployment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Jiecheng, Hu, Qinghao, Jin, Yuyang, Wang, Zerui, Sun, Peng, Gu, Yuzhe, Zhang, Wenwei, Zhai, Mingshu, Zhang, Xingcheng, Zhang, Weiming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918159092547584
author Zhou, Jiecheng
Hu, Qinghao
Jin, Yuyang
Wang, Zerui
Sun, Peng
Gu, Yuzhe
Zhang, Wenwei
Zhai, Mingshu
Zhang, Xingcheng
Zhang, Weiming
author_facet Zhou, Jiecheng
Hu, Qinghao
Jin, Yuyang
Wang, Zerui
Sun, Peng
Gu, Yuzhe
Zhang, Wenwei
Zhai, Mingshu
Zhang, Xingcheng
Zhang, Weiming
contents Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent months to enhance their reasoning and understanding abilities. However, its complex data flows and diverse tasks pose substantial challenges to RL training systems, and there is limited understanding of RLVR from a system perspective. To thoroughly understand the system challenges introduced by RLVR, we present a characterization study of RLVR tasks in our LLM deployment. Specifically, we investigate the distribution and variation trends of workloads across different RL tasks across training steps. We identify issues such as GPU idling caused by skewed sequence length distribution, inefficient parallel strategies in dynamically varying workloads, inefficient data management mechanisms, and load imbalance. We describe our observations and call for further investigation into the remaining open challenges. Furthermore, we propose PolyTrace benchmark suite to conduct evaluation with realistic workloads, and a practical use case validates that PolyTrace benchmark suite exhibits 94.7% accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25279
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RL in the Wild: Characterizing RLVR Training in LLM Deployment
Zhou, Jiecheng
Hu, Qinghao
Jin, Yuyang
Wang, Zerui
Sun, Peng
Gu, Yuzhe
Zhang, Wenwei
Zhai, Mingshu
Zhang, Xingcheng
Zhang, Weiming
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent months to enhance their reasoning and understanding abilities. However, its complex data flows and diverse tasks pose substantial challenges to RL training systems, and there is limited understanding of RLVR from a system perspective. To thoroughly understand the system challenges introduced by RLVR, we present a characterization study of RLVR tasks in our LLM deployment. Specifically, we investigate the distribution and variation trends of workloads across different RL tasks across training steps. We identify issues such as GPU idling caused by skewed sequence length distribution, inefficient parallel strategies in dynamically varying workloads, inefficient data management mechanisms, and load imbalance. We describe our observations and call for further investigation into the remaining open challenges. Furthermore, we propose PolyTrace benchmark suite to conduct evaluation with realistic workloads, and a practical use case validates that PolyTrace benchmark suite exhibits 94.7% accuracy.
title RL in the Wild: Characterizing RLVR Training in LLM Deployment
topic Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2509.25279