ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hao, Liu, Mingjie, Zhang, Shaokun, Han, Songyang, Hu, Jian, Jin, Zhenghui, Zhang, Yuchi, Diao, Shizhe, Lu, Ximing, Xu, Binfeng, Yu, Zhiding, Kautz, Jan, Dong, Yi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914409213853696
author Zhang, Hao
Liu, Mingjie
Zhang, Shaokun
Han, Songyang
Hu, Jian
Jin, Zhenghui
Zhang, Yuchi
Diao, Shizhe
Lu, Ximing
Xu, Binfeng
Yu, Zhiding
Kautz, Jan
Dong, Yi
author_facet Zhang, Hao
Liu, Mingjie
Zhang, Shaokun
Han, Songyang
Hu, Jian
Jin, Zhenghui
Zhang, Yuchi
Diao, Shizhe
Lu, Ximing
Xu, Binfeng
Yu, Zhiding
Kautz, Jan
Dong, Yi
contents Multi-turn LLM agents are increasingly important for solving complex, interactive tasks, and reinforcement learning (RL) is a key ingredient for improving their long-horizon behavior. However, RL training requires generating large numbers of sandboxed rollout trajectories, and existing infrastructures often couple rollout orchestration with the training loop, making systems hard to migrate and maintain. Under the rollout-as-a-service philosophy, we present ProRL Agent , a scalable infrastructure that serves the full agentic rollout lifecycle through an API service. ProRL Agent also provides standardized and extensible sandbox environments that support diverse agentic tasks in rootless HPC settings. We validate ProRL Agent through RL training on software engineering, math, STEM, and coding tasks. ProRL Agent is open-sourced and integrated as part of NVIDIA NeMo Gym.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18815
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
Zhang, Hao
Liu, Mingjie
Zhang, Shaokun
Han, Songyang
Hu, Jian
Jin, Zhenghui
Zhang, Yuchi
Diao, Shizhe
Lu, Ximing
Xu, Binfeng
Yu, Zhiding
Kautz, Jan
Dong, Yi
Artificial Intelligence
Multi-turn LLM agents are increasingly important for solving complex, interactive tasks, and reinforcement learning (RL) is a key ingredient for improving their long-horizon behavior. However, RL training requires generating large numbers of sandboxed rollout trajectories, and existing infrastructures often couple rollout orchestration with the training loop, making systems hard to migrate and maintain. Under the rollout-as-a-service philosophy, we present ProRL Agent , a scalable infrastructure that serves the full agentic rollout lifecycle through an API service. ProRL Agent also provides standardized and extensible sandbox environments that support diverse agentic tasks in rootless HPC settings. We validate ProRL Agent through RL training on software engineering, math, STEM, and coding tasks. ProRL Agent is open-sourced and integrated as part of NVIDIA NeMo Gym.
title ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2603.18815