SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Shiyi, Li, Dacheng, Zhao, Fangzhou, Yuan, Shuo, Hegde, Sumanth R., Chen, Connor, Ruan, Charlie, Griggs, Tyler, Liu, Shu, Tang, Eric, Liaw, Richard, Moritz, Philipp, Zaharia, Matei, Gonzalez, Joseph E., Stoica, Ion
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911277349076992
author Cao, Shiyi
Li, Dacheng
Zhao, Fangzhou
Yuan, Shuo
Hegde, Sumanth R.
Chen, Connor
Ruan, Charlie
Griggs, Tyler
Liu, Shu
Tang, Eric
Liaw, Richard
Moritz, Philipp
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
author_facet Cao, Shiyi
Li, Dacheng
Zhao, Fangzhou
Yuan, Shuo
Hegde, Sumanth R.
Chen, Connor
Ruan, Charlie
Griggs, Tyler
Liu, Shu
Tang, Eric
Liaw, Richard
Moritz, Philipp
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
contents We introduce SkyRL-Agent, a framework for efficient, multi-turn, long-horizon agent training and evaluation. It provides efficient asynchronous dispatching, lightweight tool integration, and flexible backend interoperability, enabling seamless use with existing RL frameworks such as SkyRL-train, VeRL, and Tinker. Using SkyRL-Agent, we train SA-SWE-32B, a software engineering agent trained from Qwen3-32B (24.4% Pass@1) purely with reinforcement learning. We introduce two key components: an optimized asynchronous pipeline dispatcher that achieves a 1.55x speedup over naive asynchronous batching, and a tool-enhanced training recipe leveraging an AST-based search tool to facilitate code navigation, boost rollout Pass@K, and improve training efficiency. Together, these optimizations enable SA-SWE-32B to reach 39.4% Pass@1 on SWE-Bench Verified with more than 2x cost reduction compared to prior models reaching similar performance. Despite being trained solely on SWE tasks, SA-SWE-32B generalizes effectively to other agentic tasks, including Terminal-Bench, BrowseComp-Plus, and WebArena. We further demonstrate SkyRL-Agent's extensibility through case studies on deep research, computer use, and memory agents, each trained using a different training backend.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
Cao, Shiyi
Li, Dacheng
Zhao, Fangzhou
Yuan, Shuo
Hegde, Sumanth R.
Chen, Connor
Ruan, Charlie
Griggs, Tyler
Liu, Shu
Tang, Eric
Liaw, Richard
Moritz, Philipp
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
Artificial Intelligence
We introduce SkyRL-Agent, a framework for efficient, multi-turn, long-horizon agent training and evaluation. It provides efficient asynchronous dispatching, lightweight tool integration, and flexible backend interoperability, enabling seamless use with existing RL frameworks such as SkyRL-train, VeRL, and Tinker. Using SkyRL-Agent, we train SA-SWE-32B, a software engineering agent trained from Qwen3-32B (24.4% Pass@1) purely with reinforcement learning. We introduce two key components: an optimized asynchronous pipeline dispatcher that achieves a 1.55x speedup over naive asynchronous batching, and a tool-enhanced training recipe leveraging an AST-based search tool to facilitate code navigation, boost rollout Pass@K, and improve training efficiency. Together, these optimizations enable SA-SWE-32B to reach 39.4% Pass@1 on SWE-Bench Verified with more than 2x cost reduction compared to prior models reaching similar performance. Despite being trained solely on SWE tasks, SA-SWE-32B generalizes effectively to other agentic tasks, including Terminal-Bench, BrowseComp-Plus, and WebArena. We further demonstrate SkyRL-Agent's extensibility through case studies on deep research, computer use, and memory agents, each trained using a different training backend.
title SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
topic Artificial Intelligence
url https://arxiv.org/abs/2511.16108