Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
Fuente:
arXiv
Saved in:
| Main Authors: | Khan, Azal Ahmad, Ahmed, Ammar, Fayyaz, Zeshan, Di, Sheng, Hong, Mingyi, Anwar, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Personalized Federated Learning Techniques: Empirical Analysis
by: Khan, Azal Ahmad, et al.
Published: (2024)
by: Khan, Azal Ahmad, et al.
Published: (2024)
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
IP-FL: Incentivized and Personalized Federated Learning
by: Khan, Ahmad Faraz, et al.
Published: (2023)
by: Khan, Ahmad Faraz, et al.
Published: (2023)
Safety Aware Task Planning via Large Language Models in Robotics
by: Khan, Azal Ahmad, et al.
Published: (2025)
by: Khan, Azal Ahmad, et al.
Published: (2025)
LADs: Leveraging LLMs for AI-Driven DevOps
by: Khan, Ahmad Faraz, et al.
Published: (2025)
by: Khan, Ahmad Faraz, et al.
Published: (2025)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
by: Wang, Mingyi, et al.
Published: (2026)
by: Wang, Mingyi, et al.
Published: (2026)
FLStore: Efficient Federated Learning Storage for non-training workloads
by: Khan, Ahmad Faraz, et al.
Published: (2025)
by: Khan, Ahmad Faraz, et al.
Published: (2025)
FLOSS: Federated Learning with Opt-Out and Straggler Support
by: Goetze, David J, et al.
Published: (2025)
by: Goetze, David J, et al.
Published: (2025)
Learning from Less: SINDy Surrogates in RL
by: Dixit, Aniket, et al.
Published: (2025)
by: Dixit, Aniket, et al.
Published: (2025)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
DriftXpress: Faster Drifting Models via Projected RKHS Fields
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
Action-Free Offline-to-Online RL via Discretised State Policies
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
ProToken: Token-Level Attribution for Federated Large Language Models
by: Gill, Waris, et al.
Published: (2026)
by: Gill, Waris, et al.
Published: (2026)
Adaptive Policy Synchronization for Scalable Reinforcement Learning
by: Lafuente-Mercado, Rodney
Published: (2025)
by: Lafuente-Mercado, Rodney
Published: (2025)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Scalable Policy-Based RL Algorithms for POMDPs
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
TLDR: Unsupervised Goal-Conditioned RL via Temporal Distance-Aware Representations
by: Bae, Junik, et al.
Published: (2024)
by: Bae, Junik, et al.
Published: (2024)
Structural Estimation of Markov Decision Processes in High-Dimensional State Space with Finite-Time Guarantees
by: Zeng, Siliang, et al.
Published: (2022)
by: Zeng, Siliang, et al.
Published: (2022)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Digital Twin Synchronization: Bridging the Sim-RL Agent to a Real-Time Robotic Additive Manufacturing Control
by: Ali, Matsive, et al.
Published: (2025)
by: Ali, Matsive, et al.
Published: (2025)
CRAFT: Forgetting-Aware Intervention-Based Adaptation for Continual Learning
by: Hossen, Md Anwar, et al.
Published: (2026)
by: Hossen, Md Anwar, et al.
Published: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
by: Duan, Xintong, et al.
Published: (2025)
by: Duan, Xintong, et al.
Published: (2025)
Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration
by: Le, Tuan, et al.
Published: (2025)
by: Le, Tuan, et al.
Published: (2025)
Policy Learning for Off-Dynamics RL with Deficient Support
by: Van, Linh Le Pham, et al.
Published: (2024)
by: Van, Linh Le Pham, et al.
Published: (2024)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
by: Hou, Hongru, et al.
Published: (2026)
by: Hou, Hongru, et al.
Published: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
by: Bhatia, Abhinav, et al.
Published: (2023)
by: Bhatia, Abhinav, et al.
Published: (2023)
Novel RL approach for efficient Elevator Group Control Systems
by: Vaartjes, Nathan, et al.
Published: (2025)
by: Vaartjes, Nathan, et al.
Published: (2025)
When Demonstrations Meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning
by: Zeng, Siliang, et al.
Published: (2023)
by: Zeng, Siliang, et al.
Published: (2023)
Similar Items
-
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
by: Ahmed, Ammar, et al.
Published: (2025) -
Personalized Federated Learning Techniques: Empirical Analysis
by: Khan, Azal Ahmad, et al.
Published: (2024) -
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025) -
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025) -
IP-FL: Incentivized and Personalized Federated Learning
by: Khan, Ahmad Faraz, et al.
Published: (2023)