Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Weixun, Xiong, Shaopan, Chen, Gengru, Gao, Wei, Guo, Sheng, He, Yancheng, Huang, Ju, Liu, Jiaheng, Li, Zhendong, Li, Xiaoyang, Liu, Zichen, Zhao, Haizhou, An, Dakai, Cao, Lunxi, Cao, Qiyang, Deng, Wanxi, Du, Feilei, Gu, Yiliang, Li, Jiahe, Li, Xiang, Liu, Mingjie, Luo, Yijia, Liu, Zihe, Wang, Yadao, Wang, Pei, Wu, Tianyuan, Wu, Yanan, Zhao, Yuheng, Zhao, Shuaibing, Yang, Jin, Yang, Siran, Tan, Yingshui, Yi, Huimin, Xu, Yuchi, Yuan, Yujin, Zhang, Xingyao, Qu, Lin, Su, Wenbo, Wang, Wei, Wang, Jiamang, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
by: Wang, Pei, et al.
Published: (2026)
by: Wang, Pei, et al.
Published: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
by: Li, Yanshi, et al.
Published: (2024)
by: Li, Yanshi, et al.
Published: (2024)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
by: Wu, Tianyuan, et al.
Published: (2026)
by: Wu, Tianyuan, et al.
Published: (2026)
Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
by: Long, Rujiao, et al.
Published: (2025)
by: Long, Rujiao, et al.
Published: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Bio-inspired Integrated Networking and Control for Large-Scale Swarm: A Hierarchical Co-design
by: Lin, Huan, et al.
Published: (2025)
by: Lin, Huan, et al.
Published: (2025)
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
fastbmRAG: A Fast Graph-Based RAG Framework for Efficient Processing of Large-Scale Biomedical Literature
by: Meng, Guofeng, et al.
Published: (2025)
by: Meng, Guofeng, et al.
Published: (2025)
Orchestrating Intelligence: Confidence-Aware Routing for Efficient Multi-Agent Collaboration across Multi-Scale Models
by: Wang, Jingbo, et al.
Published: (2026)
by: Wang, Jingbo, et al.
Published: (2026)
Semantic Communication via Rate Distortion Perception Bottleneck
by: Zhao, Zihe, et al.
Published: (2024)
by: Zhao, Zihe, et al.
Published: (2024)
Grams: Gradient Descent with Adaptive Momentum Scaling
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Large‐Scale, Stretchable, Self‐Protective, and Multifunctional Perovskite Luminescent Filament with Ultra‐High Stability
by: Liyan Yang, et al.
Published: (2024)
by: Liyan Yang, et al.
Published: (2024)
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
by: Deng, Yuhai, et al.
Published: (2026)
by: Deng, Yuhai, et al.
Published: (2026)
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
by: Lai, Xin, et al.
Published: (2025)
by: Lai, Xin, et al.
Published: (2025)
A Robust Multi-Scale Framework with Test-Time Adaptation for sEEG-Based Speech Decoding
by: Wang, Suli, et al.
Published: (2025)
by: Wang, Suli, et al.
Published: (2025)
Comparative diagnostic accuracy of next‐generation sequencing in different specimen types for periprosthetic joint infection: A systematic review and meta‐analysis
by: Lina Wang, et al.
Published: (2025)
by: Lina Wang, et al.
Published: (2025)
Physiologically Based Pharmacokinetic Modeling to Investigate the Disease‐Drug–Drug Interactions between Voriconazole and Nirmatrelvir/Ritonavir in COVID‐19 Patients with CYP2C19 Phenotypes
by: Peile Wang, et al.
Published: (2024)
by: Peile Wang, et al.
Published: (2024)
Fuzzy containment control for multi‐USV systems based on communication topology reconstruction
by: Haoran Zhao, et al.
Published: (2026)
by: Haoran Zhao, et al.
Published: (2026)
Large‐Scale Synthesis of High‐Purity Isoguanosine and Resolution of its Crystal Structure by Microcrystal Electron Diffraction
by: Kaichao Wang, et al.
Published: (2024)
by: Kaichao Wang, et al.
Published: (2024)
A Federated Learning Framework for Handling Subtype Confounding and Heterogeneity in Large-Scale Neuroimaging Diagnosis
by: Zhao, Xinglin, et al.
Published: (2025)
by: Zhao, Xinglin, et al.
Published: (2025)
Constructing Ag/Cu 2 O Interface for Efficient Neutral CO 2 Electroreduction to C 2 H 4
by: Zongnan Wei, et al.
Published: (2024)
by: Zongnan Wei, et al.
Published: (2024)
Constructing Ag/Cu 2 O Interface for Efficient Neutral CO 2 Electroreduction to C 2 H 4
by: Zongnan Wei, et al.
Published: (2024)
by: Zongnan Wei, et al.
Published: (2024)
Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation
by: He, Xiao, et al.
Published: (2026)
by: He, Xiao, et al.
Published: (2026)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Controlling the Latent Diffusion Model for Generative Image Shadow Removal via Residual Generation
by: Li, Xinjie, et al.
Published: (2024)
by: Li, Xinjie, et al.
Published: (2024)
Mechanistic Impacts of a Scale‐Aware Convection Scheme on Typhoon Intensity: Diagnostics From Minimum Sea‐Level Pressure
by: Yanjie Liu, et al.
Published: (2025)
by: Yanjie Liu, et al.
Published: (2025)
Private Private Information in Second-Price Auction
by: Liu, Boyu, et al.
Published: (2026)
by: Liu, Boyu, et al.
Published: (2026)
Finasteride Use Does Not Lead to Depression or Suicide: Insights From a Large‐Scale Cohort Study and Mendelian Randomization Analysis
by: Jing Wang, et al.
Published: (2025)
by: Jing Wang, et al.
Published: (2025)
Tensile Deformation Law and Scale Effect of Low‐Constraint Ultra‐Thin Fiber Metal Laminates
by: Yanfeng Zhang, et al.
Published: (2025)
by: Yanfeng Zhang, et al.
Published: (2025)
PhiNet: Speaker Verification with Phonetic Interpretability
by: Ma, Yi, et al.
Published: (2026)
by: Ma, Yi, et al.
Published: (2026)
Similar Items
-
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025) -
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025) -
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026) -
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025) -
ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
by: Wang, Pei, et al.
Published: (2026)