On-Policy RL with Optimal Reward Baseline

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Yaru, Dong, Li, Wu, Xun, Huang, Shaohan, Chi, Zewen, Wei, Furu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909636456611840
author Hao, Yaru
Dong, Li
Wu, Xun
Huang, Shaohan
Chi, Zewen
Wei, Furu
author_facet Hao, Yaru
Dong, Li
Wu, Xun
Huang, Shaohan
Chi, Zewen
Wei, Furu
contents Reinforcement learning algorithms are fundamental to align large language models with human preferences and to enhance their reasoning capabilities. However, current reinforcement learning algorithms often suffer from training instability due to loose on-policy constraints and computational inefficiency due to auxiliary models. In this work, we propose On-Policy RL with Optimal reward baseline (OPO), a novel and simplified reinforcement learning algorithm designed to address these challenges. OPO emphasizes the importance of exact on-policy training, which empirically stabilizes the training process and enhances exploration. Moreover, OPO integrates a practically feasible formulation of the optimal reward baseline that minimizes gradient variance. We evaluate OPO on mathematical reasoning benchmarks. The results demonstrate its superior performance and training stability without additional models or regularization terms. Furthermore, OPO achieves lower policy shifts and higher output entropy, encouraging more diverse and less repetitive responses. These results highlight OPO as a promising direction for stable and effective reinforcement learning in large language model alignment and reasoning tasks. The implementation is merged into the verl library at https://verl.readthedocs.io/en/latest/algo/opo.html.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On-Policy RL with Optimal Reward Baseline
Hao, Yaru
Dong, Li
Wu, Xun
Huang, Shaohan
Chi, Zewen
Wei, Furu
Machine Learning
Computation and Language
Reinforcement learning algorithms are fundamental to align large language models with human preferences and to enhance their reasoning capabilities. However, current reinforcement learning algorithms often suffer from training instability due to loose on-policy constraints and computational inefficiency due to auxiliary models. In this work, we propose On-Policy RL with Optimal reward baseline (OPO), a novel and simplified reinforcement learning algorithm designed to address these challenges. OPO emphasizes the importance of exact on-policy training, which empirically stabilizes the training process and enhances exploration. Moreover, OPO integrates a practically feasible formulation of the optimal reward baseline that minimizes gradient variance. We evaluate OPO on mathematical reasoning benchmarks. The results demonstrate its superior performance and training stability without additional models or regularization terms. Furthermore, OPO achieves lower policy shifts and higher output entropy, encouraging more diverse and less repetitive responses. These results highlight OPO as a promising direction for stable and effective reinforcement learning in large language model alignment and reasoning tasks. The implementation is merged into the verl library at https://verl.readthedocs.io/en/latest/algo/opo.html.
title On-Policy RL with Optimal Reward Baseline
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.23585