SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Shiyi, Li, Dacheng, Zhao, Fangzhou, Yuan, Shuo, Hegde, Sumanth R., Chen, Connor, Ruan, Charlie, Griggs, Tyler, Liu, Shu, Tang, Eric, Liaw, Richard, Moritz, Philipp, Zaharia, Matei, Gonzalez, Joseph E., Stoica, Ion |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
by: Griggs, Tyler, et al.
Published: (2026)
by: Griggs, Tyler, et al.
Published: (2026)
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
by: Davis, Jared Quincy, et al.
Published: (2024)
by: Davis, Jared Quincy, et al.
Published: (2024)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
by: Yang, Haojin, et al.
Published: (2026)
by: Yang, Haojin, et al.
Published: (2026)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
Optimizing LLM Queries in Relational Data Analytics Workloads
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
by: Zhao, Weikang, et al.
Published: (2025)
by: Zhao, Weikang, et al.
Published: (2025)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
by: Gao, Jiaxuan, et al.
Published: (2026)
by: Gao, Jiaxuan, et al.
Published: (2026)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
AI-Driven Research for Databases
by: Cheng, Audrey, et al.
Published: (2026)
by: Cheng, Audrey, et al.
Published: (2026)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
Training RL Agents for Multi-Objective Network Defense Tasks
by: Molina-Markham, Andres, et al.
Published: (2025)
by: Molina-Markham, Andres, et al.
Published: (2025)
OpenClaw-RL: Train Any Agent Simply by Talking
by: Wang, Yinjie, et al.
Published: (2026)
by: Wang, Yinjie, et al.
Published: (2026)
Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
by: Mai, Xinji, et al.
Published: (2025)
by: Mai, Xinji, et al.
Published: (2025)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Accelerating Direct Preference Optimization with Prefix Sharing
by: Wang, Franklin, et al.
Published: (2024)
by: Wang, Franklin, et al.
Published: (2024)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026)
by: Pan, Melissa Z., et al.
Published: (2026)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
by: Cao, Shiyi, et al.
Published: (2026)
by: Cao, Shiyi, et al.
Published: (2026)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
by: Bai, Hao, et al.
Published: (2024)
by: Bai, Hao, et al.
Published: (2024)
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
by: Mehta, Sushant, et al.
Published: (2026)
by: Mehta, Sushant, et al.
Published: (2026)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
by: Zhou, Yifei, et al.
Published: (2025)
by: Zhou, Yifei, et al.
Published: (2025)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Similar Items
-
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
by: Li, Dacheng, et al.
Published: (2025) -
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024) -
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025) -
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024) -
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)