Bounded Ratio Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ao, Yunke, Chen, Le, Lee, Bruce D., Wahd, Assefa S., Czarnobai, Aline, Fürnstahl, Philipp, Schölkopf, Bernhard, Krause, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910179184869376
author Ao, Yunke
Chen, Le
Lee, Bruce D.
Wahd, Assefa S.
Czarnobai, Aline
Fürnstahl, Philipp
Schölkopf, Bernhard
Krause, Andreas
author_facet Ao, Yunke
Chen, Le
Lee, Bruce D.
Wahd, Assefa S.
Czarnobai, Aline
Fürnstahl, Philipp
Schölkopf, Bernhard
Krause, Andreas
contents Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect between the underlying foundations of trust region methods and the heuristic clipped objective used in PPO. In this paper, we bridge this gap by introducing the Bounded Ratio Reinforcement Learning (BRRL) framework. We formulate a novel regularized and constrained policy optimization problem and derive its analytical optimal solution. We prove that this solution ensures monotonic performance improvement. To handle parameterized policy classes, we develop a policy optimization algorithm called Bounded Policy Optimization (BPO) that minimizes an advantage-weighted divergence between the policy and the analytic optimal solution from BRRL. We further establish a lower bound on the expected performance of the resulting policy in terms of the BPO loss function. Notably, our framework also provides a new theoretical lens to interpret the success of the PPO loss, and connects trust region policy optimization and the Cross-Entropy Method (CEM). We additionally extend BPO to Group-relative BPO (GBPO) for LLM fine-tuning. Empirical evaluations of BPO across MuJoCo, Atari, and complex IsaacLab environments (e.g., Humanoid locomotion), and of GBPO for LLM fine-tuning tasks, demonstrate that BPO and GBPO generally match or outperform PPO and GRPO in stability and final performance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18578
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bounded Ratio Reinforcement Learning
Ao, Yunke
Chen, Le
Lee, Bruce D.
Wahd, Assefa S.
Czarnobai, Aline
Fürnstahl, Philipp
Schölkopf, Bernhard
Krause, Andreas
Machine Learning
Artificial Intelligence
I.2.6
Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect between the underlying foundations of trust region methods and the heuristic clipped objective used in PPO. In this paper, we bridge this gap by introducing the Bounded Ratio Reinforcement Learning (BRRL) framework. We formulate a novel regularized and constrained policy optimization problem and derive its analytical optimal solution. We prove that this solution ensures monotonic performance improvement. To handle parameterized policy classes, we develop a policy optimization algorithm called Bounded Policy Optimization (BPO) that minimizes an advantage-weighted divergence between the policy and the analytic optimal solution from BRRL. We further establish a lower bound on the expected performance of the resulting policy in terms of the BPO loss function. Notably, our framework also provides a new theoretical lens to interpret the success of the PPO loss, and connects trust region policy optimization and the Cross-Entropy Method (CEM). We additionally extend BPO to Group-relative BPO (GBPO) for LLM fine-tuning. Empirical evaluations of BPO across MuJoCo, Atari, and complex IsaacLab environments (e.g., Humanoid locomotion), and of GBPO for LLM fine-tuning tasks, demonstrate that BPO and GBPO generally match or outperform PPO and GRPO in stability and final performance.
title Bounded Ratio Reinforcement Learning
topic Machine Learning
Artificial Intelligence
I.2.6
url https://arxiv.org/abs/2604.18578