AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hanchen, Liu, Xiao, Lv, Bowen, Sun, Xueqiao, Jing, Bohao, Iong, Iat Long, Hou, Zhenyu, Qi, Zehan, Lai, Hanyu, Xu, Yifan, Lu, Rui, Wang, Hongning, Tang, Jie, Dong, Yuxiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915533810565120
author Zhang, Hanchen
Liu, Xiao
Lv, Bowen
Sun, Xueqiao
Jing, Bohao
Iong, Iat Long
Hou, Zhenyu
Qi, Zehan
Lai, Hanyu
Xu, Yifan
Lu, Rui
Wang, Hongning
Tang, Jie
Dong, Yuxiao
author_facet Zhang, Hanchen
Liu, Xiao
Lv, Bowen
Sun, Xueqiao
Jing, Bohao
Iong, Iat Long
Hou, Zhenyu
Qi, Zehan
Lai, Hanyu
Xu, Yifan
Lu, Rui
Wang, Hongning
Tang, Jie
Dong, Yuxiao
contents Recent advances in large language models (LLMs) have sparked growing interest in building generalist agents that can learn through online interactions. However, applying reinforcement learning (RL) to train LLM agents in multi-turn, multi-task settings remains challenging due to lack of scalable infrastructure and stable training algorithms. In this work, we present the AgentRL framework for scalable multi-turn, multi-task agentic RL training. On the infrastructure side, AgentRL features a fully-asynchronous generation-training pipeline for efficient multi-turn RL. To support heterogeneous environment development in multi-task RL, we design a unified function-call based API interface, containerized environment development, and a centralized controller. On the algorithm side, we propose cross-policy sampling to encourage model exploration in multi-turn settings and task advantage normalization to stabilize multi-task training. Experiments show that AgentRL, trained on open LLMs across five agentic tasks, significantly outperforms GPT-5, Clause-Sonnet-4, DeepSeek-R1, and other open-source LLM agents. Multi-task training with AgentRL matches the best results among all task-specific models. AgentRL is open-sourced at https://github.com/THUDM/AgentRL. The algorithm and framework are adopted in building \textsc{\href{https://autoglm.zhipuai.cn}{AutoGLM}}.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04206
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
Zhang, Hanchen
Liu, Xiao
Lv, Bowen
Sun, Xueqiao
Jing, Bohao
Iong, Iat Long
Hou, Zhenyu
Qi, Zehan
Lai, Hanyu
Xu, Yifan
Lu, Rui
Wang, Hongning
Tang, Jie
Dong, Yuxiao
Artificial Intelligence
Recent advances in large language models (LLMs) have sparked growing interest in building generalist agents that can learn through online interactions. However, applying reinforcement learning (RL) to train LLM agents in multi-turn, multi-task settings remains challenging due to lack of scalable infrastructure and stable training algorithms. In this work, we present the AgentRL framework for scalable multi-turn, multi-task agentic RL training. On the infrastructure side, AgentRL features a fully-asynchronous generation-training pipeline for efficient multi-turn RL. To support heterogeneous environment development in multi-task RL, we design a unified function-call based API interface, containerized environment development, and a centralized controller. On the algorithm side, we propose cross-policy sampling to encourage model exploration in multi-turn settings and task advantage normalization to stabilize multi-task training. Experiments show that AgentRL, trained on open LLMs across five agentic tasks, significantly outperforms GPT-5, Clause-Sonnet-4, DeepSeek-R1, and other open-source LLM agents. Multi-task training with AgentRL matches the best results among all task-specific models. AgentRL is open-sourced at https://github.com/THUDM/AgentRL. The algorithm and framework are adopted in building \textsc{\href{https://autoglm.zhipuai.cn}{AutoGLM}}.
title AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
topic Artificial Intelligence
url https://arxiv.org/abs/2510.04206