Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fakoor, Rasool, Aubry, Murdock, Stranges, Nicholas, Smola, Alexander J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
Trust-Region Adaptive Policy Optimization
von: Su, Mingyu, et al.
Veröffentlicht: (2025)
von: Su, Mingyu, et al.
Veröffentlicht: (2025)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
Policy Learning for Off-Dynamics RL with Deficient Support
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
von: Lee, Deokjae, et al.
Veröffentlicht: (2024)
von: Lee, Deokjae, et al.
Veröffentlicht: (2024)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
Scalable On-Policy Reinforcement Learning via Adaptive Batch Scaling
von: Park, Jongchan
Veröffentlicht: (2026)
von: Park, Jongchan
Veröffentlicht: (2026)
Learning the Target Network in Function Space
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
von: Nie, Dong
Veröffentlicht: (2026)
von: Nie, Dong
Veröffentlicht: (2026)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
Matrix Low-Rank Trust Region Policy Optimization
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
von: Guan, Zhong, et al.
Veröffentlicht: (2026)
von: Guan, Zhong, et al.
Veröffentlicht: (2026)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
Transformer Block Coupling and its Correlation with Generalization in LLMs
von: Aubry, Murdock, et al.
Veröffentlicht: (2024)
von: Aubry, Murdock, et al.
Veröffentlicht: (2024)
Proximal Policy Optimization with Adaptive Exploration
von: Lixandru, Andrei
Veröffentlicht: (2024)
von: Lixandru, Andrei
Veröffentlicht: (2024)
APC-RL: Exceeding Data-Driven Behavior Priors with Adaptive Policy Composition
von: Rietz, Finn, et al.
Veröffentlicht: (2026)
von: Rietz, Finn, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025) -
Trust-Region Adaptive Policy Optimization
von: Su, Mingyu, et al.
Veröffentlicht: (2025) -
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026) -
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)