RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Wei, Zhao, Yuheng, Wu, Tianyuan, Xiong, Shaopan, Wang, Weixun, An, Dakai, Cao, Lunxi, Muhtar, Dilxat, Liu, Zichen, Zhao, Haizhou, Huang, Ju, Yang, Siran, Li, Yongbin, Su, Wenbo, Wang, Jiamang, Qu, Lin, Zheng, Bo, Wang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
by: Wang, Weixun, et al.
Published: (2025)
by: Wang, Weixun, et al.
Published: (2025)
Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
by: Wang, Weixun, et al.
Published: (2025)
by: Wang, Weixun, et al.
Published: (2025)
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
by: Wu, Tianyuan, et al.
Published: (2026)
by: Wu, Tianyuan, et al.
Published: (2026)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Remote Sensing Image Super-Resolution for Imbalanced Textures: A Texture-Aware Diffusion Framework
by: Zhang, Enzhuo, et al.
Published: (2026)
by: Zhang, Enzhuo, et al.
Published: (2026)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
by: Wang, Pei, et al.
Published: (2026)
by: Wang, Pei, et al.
Published: (2026)
LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
by: Muhtar, Dilxat, et al.
Published: (2024)
by: Muhtar, Dilxat, et al.
Published: (2024)
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
by: Li, Pengxiang, et al.
Published: (2026)
by: Li, Pengxiang, et al.
Published: (2026)
Carbon Nitrides‐Based Heterojunction for High‐Efficient Li Salt Dissociation
by: Minchen Hou, et al.
Published: (2025)
by: Minchen Hou, et al.
Published: (2025)
AMAP Agentic Planning Technical Report
by: AMAP AI Agent Team, et al.
Published: (2025)
by: AMAP AI Agent Team, et al.
Published: (2025)
GEM: A Gym for Agentic LLMs
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
by: Li, Zhenshi, et al.
Published: (2024)
by: Li, Zhenshi, et al.
Published: (2024)
Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models
by: Muhtar, Dilxat, et al.
Published: (2025)
by: Muhtar, Dilxat, et al.
Published: (2025)
When Does Sparsity Mitigate the Curse of Depth in LLMs
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
PyVision-RL: Forging Open Agentic Vision Models via RL
by: Zhao, Shitian, et al.
Published: (2026)
by: Zhao, Shitian, et al.
Published: (2026)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
by: Wang, Xiaoyin, et al.
Published: (2024)
by: Wang, Xiaoyin, et al.
Published: (2024)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
by: Tan, Xin, et al.
Published: (2026)
by: Tan, Xin, et al.
Published: (2026)
FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding
by: Li, Zhenshi, et al.
Published: (2025)
by: Li, Zhenshi, et al.
Published: (2025)
BREAKING THE SILENCE OF THE LAMBS: INTEGRATING MEDICAL STAFF IN PREVENTION OF HUMAN TRAFFICKING
by: Muhtar Cokar
Published: (2016)
by: Muhtar Cokar
Published: (2016)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Multidimensional Dysfunction in Chronic Nonspecific Low Back Pain: A Correlational Study of Key Clinical Measures
by: Zhao Wang, et al.
Published: (2026)
by: Zhao Wang, et al.
Published: (2026)
Diffusion Language Models Know the Answer Before Decoding
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
Skill Reuse as Compression in Agentic RL
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
RAGEN-2: Reasoning Collapse in Agentic RL
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
A Comprehensive PPG-based Dataset for HR/HRV Studies
by: Xu, Jingye, et al.
Published: (2025)
by: Xu, Jingye, et al.
Published: (2025)
ProgCo: Program Helps Self-Correction of Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2025)
by: Song, Xiaoshuai, et al.
Published: (2025)
Co‐Doping Engineered High Performance Ni‐Rich Layered Cathode
by: Kaili Li, et al.
Published: (2025)
by: Kaili Li, et al.
Published: (2025)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Similar Items
-
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025) -
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026) -
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025) -
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026) -
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)