Soft Policy Optimization: Online Off-Policy RL for Sequence Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Taco, Zhang, David W., Zheng, Kunhao, Tang, Yunhao, Munos, Remi, Synnaeve, Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915186239078400
author Cohen, Taco
Zhang, David W.
Zheng, Kunhao
Tang, Yunhao
Munos, Remi
Synnaeve, Gabriel
author_facet Cohen, Taco
Zhang, David W.
Zheng, Kunhao
Tang, Yunhao
Munos, Remi
Synnaeve, Gabriel
contents RL-based post-training of language models is almost exclusively done using on-policy methods such as PPO. These methods cannot learn from arbitrary sequences such as those produced earlier in training, in earlier runs, by human experts or other policies, or by decoding and exploration methods. This results in severe sample inefficiency and exploration difficulties, as well as a potential loss of diversity in the policy responses. Moreover, asynchronous PPO implementations require frequent and costly model transfers, and typically use value models which require a large amount of memory. In this paper we introduce Soft Policy Optimization (SPO), a simple, scalable and principled Soft RL method for sequence model policies that can learn from arbitrary online and offline trajectories and does not require a separate value model. In experiments on code contests, we shows that SPO outperforms PPO on pass@10, is significantly faster and more memory efficient, is able to benefit from off-policy data, enjoys improved stability, and learns more diverse (i.e. soft) policies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Soft Policy Optimization: Online Off-Policy RL for Sequence Models
Cohen, Taco
Zhang, David W.
Zheng, Kunhao
Tang, Yunhao
Munos, Remi
Synnaeve, Gabriel
Machine Learning
Artificial Intelligence
RL-based post-training of language models is almost exclusively done using on-policy methods such as PPO. These methods cannot learn from arbitrary sequences such as those produced earlier in training, in earlier runs, by human experts or other policies, or by decoding and exploration methods. This results in severe sample inefficiency and exploration difficulties, as well as a potential loss of diversity in the policy responses. Moreover, asynchronous PPO implementations require frequent and costly model transfers, and typically use value models which require a large amount of memory. In this paper we introduce Soft Policy Optimization (SPO), a simple, scalable and principled Soft RL method for sequence model policies that can learn from arbitrary online and offline trajectories and does not require a separate value model. In experiments on code contests, we shows that SPO outperforms PPO on pass@10, is significantly faster and more memory efficient, is able to benefit from off-policy data, enjoys improved stability, and learns more diverse (i.e. soft) policies.
title Soft Policy Optimization: Online Off-Policy RL for Sequence Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.05453