EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lunjun, Ba, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
by: Block, Adam, et al.
Published: (2025)
by: Block, Adam, et al.
Published: (2025)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
by: Xu, Cong, et al.
Published: (2025)
by: Xu, Cong, et al.
Published: (2025)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
Tacit Learning with Adaptive Information Selection for Cooperative Multi-Agent Reinforcement Learning
by: Liu, Lunjun, et al.
Published: (2024)
by: Liu, Lunjun, et al.
Published: (2024)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
FedEMA-Distill: Exponential Moving Average Guided Knowledge Distillation for Robust Federated Learning
by: Reguieg, Hamza, et al.
Published: (2026)
by: Reguieg, Hamza, et al.
Published: (2026)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Kimi k1.5: Scaling Reinforcement Learning with LLMs
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
by: Lee, Taeho, et al.
Published: (2026)
by: Lee, Taeho, et al.
Published: (2026)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Reinforcement Learning with Rubric Anchors
by: Huang, Zenan, et al.
Published: (2025)
by: Huang, Zenan, et al.
Published: (2025)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
by: Batra, Sumeet, et al.
Published: (2023)
by: Batra, Sumeet, et al.
Published: (2023)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
by: Lehmann, Matthias
Published: (2024)
by: Lehmann, Matthias
Published: (2024)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
Mastering Diverse Domains through World Models
by: Hafner, Danijar, et al.
Published: (2023)
by: Hafner, Danijar, et al.
Published: (2023)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
by: Zhou, Yuzhen, et al.
Published: (2025)
by: Zhou, Yuzhen, et al.
Published: (2025)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
by: Huo, Yingxiao, et al.
Published: (2026)
by: Huo, Yingxiao, et al.
Published: (2026)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
by: Zhang, Yixian, et al.
Published: (2025)
by: Zhang, Yixian, et al.
Published: (2025)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
by: Hou, Hongru, et al.
Published: (2026)
by: Hou, Hongru, et al.
Published: (2026)
Using Large Language Models for Hyperparameter Optimization
by: Zhang, Michael R., et al.
Published: (2023)
by: Zhang, Michael R., et al.
Published: (2023)
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
by: Schoepp, Sheila, et al.
Published: (2024)
by: Schoepp, Sheila, et al.
Published: (2024)
Similar Items
-
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
by: Block, Adam, et al.
Published: (2025) -
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
by: Zhang, Lunjun, et al.
Published: (2026) -
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
by: Xu, Cong, et al.
Published: (2025) -
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025) -
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)