Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Luke J., Zhang, Zhuoyang, Hu, Qinghao, Yang, Shang, Han, Song |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hide to Guide: Learning via Semantic Masking
por: Liu, Ruitao, et al.
Publicado: (2026)
por: Liu, Ruitao, et al.
Publicado: (2026)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
por: Ye, Chenlu, et al.
Publicado: (2026)
por: Ye, Chenlu, et al.
Publicado: (2026)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
por: Cohen, Taco, et al.
Publicado: (2025)
por: Cohen, Taco, et al.
Publicado: (2025)
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
por: Lu, Jian
Publicado: (2025)
por: Lu, Jian
Publicado: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
por: Fakoor, Rasool, et al.
Publicado: (2026)
por: Fakoor, Rasool, et al.
Publicado: (2026)
Policy Learning for Off-Dynamics RL with Deficient Support
por: Van, Linh Le Pham, et al.
Publicado: (2024)
por: Van, Linh Le Pham, et al.
Publicado: (2024)
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
por: Zheng, Haizhong, et al.
Publicado: (2025)
por: Zheng, Haizhong, et al.
Publicado: (2025)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
por: Lee, Haanvid, et al.
Publicado: (2024)
por: Lee, Haanvid, et al.
Publicado: (2024)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
por: Goodall, Alexander W., et al.
Publicado: (2025)
por: Goodall, Alexander W., et al.
Publicado: (2025)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
por: Hu, Qinghao, et al.
Publicado: (2025)
por: Hu, Qinghao, et al.
Publicado: (2025)
Partial Policy Gradients for RL in LLMs
por: Mathur, Puneet, et al.
Publicado: (2026)
por: Mathur, Puneet, et al.
Publicado: (2026)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
por: Romeo, Carlo, et al.
Publicado: (2026)
por: Romeo, Carlo, et al.
Publicado: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025)
por: Lu, Han, et al.
Publicado: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
por: Shen, Guobin, et al.
Publicado: (2026)
por: Shen, Guobin, et al.
Publicado: (2026)
Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
por: Gu, Yuxian, et al.
Publicado: (2025)
por: Gu, Yuxian, et al.
Publicado: (2025)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
por: Guan, Zhong, et al.
Publicado: (2026)
por: Guan, Zhong, et al.
Publicado: (2026)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
por: Zhang, Wenhao, et al.
Publicado: (2025)
por: Zhang, Wenhao, et al.
Publicado: (2025)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
por: Luo, Yu, et al.
Publicado: (2026)
por: Luo, Yu, et al.
Publicado: (2026)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)
por: Noukhovitch, Michael, et al.
Publicado: (2024)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
por: Kiyohara, Haruka, et al.
Publicado: (2023)
por: Kiyohara, Haruka, et al.
Publicado: (2023)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
por: Luo, Yu, et al.
Publicado: (2024)
por: Luo, Yu, et al.
Publicado: (2024)
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
por: Hu, Zengjie, et al.
Publicado: (2025)
por: Hu, Zengjie, et al.
Publicado: (2025)
EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
por: Zhang, Zhuoyang, et al.
Publicado: (2024)
por: Zhang, Zhuoyang, et al.
Publicado: (2024)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
por: Shah, Jay, et al.
Publicado: (2024)
por: Shah, Jay, et al.
Publicado: (2024)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
por: Zhang, Kaichen, et al.
Publicado: (2025)
por: Zhang, Kaichen, et al.
Publicado: (2025)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
por: Sun, Cheng, et al.
Publicado: (2025)
por: Sun, Cheng, et al.
Publicado: (2025)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
por: Li, Yingru, et al.
Publicado: (2026)
por: Li, Yingru, et al.
Publicado: (2026)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
por: Zhang, Yiqi, et al.
Publicado: (2026)
por: Zhang, Yiqi, et al.
Publicado: (2026)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
por: Meng, Wenjia, et al.
Publicado: (2024)
por: Meng, Wenjia, et al.
Publicado: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
por: Xi, Zhiheng, et al.
Publicado: (2025)
por: Xi, Zhiheng, et al.
Publicado: (2025)
On Entropy Control in LLM-RL Algorithms
por: Shen, Han
Publicado: (2025)
por: Shen, Han
Publicado: (2025)
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
por: Zhang, Zhuoyang, et al.
Publicado: (2026)
por: Zhang, Zhuoyang, et al.
Publicado: (2026)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
por: Cai, Xin-Qiang, et al.
Publicado: (2026)
Low Variance Off-policy Evaluation with State-based Importance Sampling
por: Bossens, David M., et al.
Publicado: (2022)
por: Bossens, David M., et al.
Publicado: (2022)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
por: Hu, Xixi, et al.
Publicado: (2024)
por: Hu, Xixi, et al.
Publicado: (2024)
Deep RL With Information Constrained Policies: Generalization in Continuous Control
por: Malloy, Tailia, et al.
Publicado: (2020)
por: Malloy, Tailia, et al.
Publicado: (2020)
SHIP: A Shapelet-based Approach for Interpretable Patient-Ventilator Asynchrony Detection
por: Le, Xuan-May, et al.
Publicado: (2025)
por: Le, Xuan-May, et al.
Publicado: (2025)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
por: Gu, Hao, et al.
Publicado: (2026)
por: Gu, Hao, et al.
Publicado: (2026)
Ejemplares similares
-
Hide to Guide: Learning via Semantic Masking
por: Liu, Ruitao, et al.
Publicado: (2026) -
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
por: Ye, Chenlu, et al.
Publicado: (2026) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
por: Cohen, Taco, et al.
Publicado: (2025) -
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
por: Lu, Jian
Publicado: (2025) -
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
por: Fakoor, Rasool, et al.
Publicado: (2026)