Stabilizing Reinforcement Learning with LLMs: Formulation and Practices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Chujie, Dang, Kai, Yu, Bowen, Li, Mingze, Jiang, Huiqiang, Lin, Junrong, Liu, Yuqiong, Lin, Hao, Wu, Chencan, Hu, Feng, Yang, An, Zhou, Jingren, Lin, Junyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909941071085568
author Zheng, Chujie
Dang, Kai
Yu, Bowen
Li, Mingze
Jiang, Huiqiang
Lin, Junrong
Liu, Yuqiong
Lin, Hao
Wu, Chencan
Hu, Feng
Yang, An
Zhou, Jingren
Lin, Junyang
author_facet Zheng, Chujie
Dang, Kai
Yu, Bowen
Li, Mingze
Jiang, Huiqiang
Lin, Junrong
Liu, Yuqiong
Lin, Hao
Wu, Chencan
Hu, Feng
Yang, An
Zhou, Jingren
Lin, Junyang
contents This paper proposes a novel formulation for reinforcement learning (RL) with large language models, explaining why and under what conditions the true sequence-level reward can be optimized via a surrogate token-level objective in policy gradient methods such as REINFORCE. Specifically, through a first-order approximation, we show that this surrogate becomes increasingly valid only when both the training-inference discrepancy and policy staleness are minimized. This insight provides a principled explanation for the crucial role of several widely adopted techniques in stabilizing RL training, including importance sampling correction, clipping, and particularly Routing Replay for Mixture-of-Experts (MoE) models. Through extensive experiments with a 30B MoE model totaling hundreds of thousands of GPU hours, we show that for on-policy training, the basic policy gradient algorithm with importance sampling correction achieves the highest training stability. When off-policy updates are introduced to accelerate convergence, combining clipping and Routing Replay becomes essential to mitigate the instability caused by policy staleness. Notably, once training is stabilized, prolonged optimization consistently yields comparable final performance regardless of cold-start initialization. We hope that the shared insights and the developed recipes for stable RL training will facilitate future research.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Zheng, Chujie
Dang, Kai
Yu, Bowen
Li, Mingze
Jiang, Huiqiang
Lin, Junrong
Liu, Yuqiong
Lin, Hao
Wu, Chencan
Hu, Feng
Yang, An
Zhou, Jingren
Lin, Junyang
Machine Learning
Artificial Intelligence
Computation and Language
This paper proposes a novel formulation for reinforcement learning (RL) with large language models, explaining why and under what conditions the true sequence-level reward can be optimized via a surrogate token-level objective in policy gradient methods such as REINFORCE. Specifically, through a first-order approximation, we show that this surrogate becomes increasingly valid only when both the training-inference discrepancy and policy staleness are minimized. This insight provides a principled explanation for the crucial role of several widely adopted techniques in stabilizing RL training, including importance sampling correction, clipping, and particularly Routing Replay for Mixture-of-Experts (MoE) models. Through extensive experiments with a 30B MoE model totaling hundreds of thousands of GPU hours, we show that for on-policy training, the basic policy gradient algorithm with importance sampling correction achieves the highest training stability. When off-policy updates are introduced to accelerate convergence, combining clipping and Routing Replay becomes essential to mitigate the instability caused by policy staleness. Notably, once training is stabilized, prolonged optimization consistently yields comparable final performance regardless of cold-start initialization. We hope that the shared insights and the developed recipes for stable RL training will facilitate future research.
title Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.01374