Non-Asymptotic Global Convergence of PPO-Clip

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yin, Dai, Qiming, Zhang, Junyu, Wen, Zaiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915684662902784
author Liu, Yin
Dai, Qiming
Zhang, Junyu
Wen, Zaiwen
author_facet Liu, Yin
Dai, Qiming
Zhang, Junyu
Wen, Zaiwen
contents Reinforcement learning (RL) has gained attention for aligning large language models (LLMs) via reinforcement learning from human feedback (RLHF). The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their efficiency. These algorithms incorporate a clipping mechanism to improve stability. Besides, a regularization term, such as the reverse KL-divergence or a more general \(f\)-divergence, is introduced to prevent policy drift. Despite their empirical success, a rigorous theoretical understanding of the problem and the algorithm's properties is limited. This paper advances the theoretical foundations of the PPO-Clip algorithm by analyzing a deterministic actor-only PPO algorithm within the general RL setting with \(f\)-divergence regularization under the softmax policy parameterization. We derive a non-uniform Lipschitz smoothness condition and a Łojasiewicz inequality for the considered problem. Based on these, a non-asymptotic linear convergence rate to the globally optimal policy is established for the forward KL-regularizer. Furthermore, stationary convergence and local linear convergence are derived for the reverse KL-regularizer.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Non-Asymptotic Global Convergence of PPO-Clip
Liu, Yin
Dai, Qiming
Zhang, Junyu
Wen, Zaiwen
Optimization and Control
Machine Learning
Reinforcement learning (RL) has gained attention for aligning large language models (LLMs) via reinforcement learning from human feedback (RLHF). The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their efficiency. These algorithms incorporate a clipping mechanism to improve stability. Besides, a regularization term, such as the reverse KL-divergence or a more general \(f\)-divergence, is introduced to prevent policy drift. Despite their empirical success, a rigorous theoretical understanding of the problem and the algorithm's properties is limited. This paper advances the theoretical foundations of the PPO-Clip algorithm by analyzing a deterministic actor-only PPO algorithm within the general RL setting with \(f\)-divergence regularization under the softmax policy parameterization. We derive a non-uniform Lipschitz smoothness condition and a Łojasiewicz inequality for the considered problem. Based on these, a non-asymptotic linear convergence rate to the globally optimal policy is established for the forward KL-regularizer. Furthermore, stationary convergence and local linear convergence are derived for the reverse KL-regularizer.
title Non-Asymptotic Global Convergence of PPO-Clip
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2512.16565