Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Arnal, Charles, Narozniak, Gaëtan, Cabannes, Vivien, Tang, Yunhao, Kempe, Julia, Munos, Remi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915643707621376
author Arnal, Charles
Narozniak, Gaëtan
Cabannes, Vivien
Tang, Yunhao
Kempe, Julia
Munos, Remi
author_facet Arnal, Charles
Narozniak, Gaëtan
Cabannes, Vivien
Tang, Yunhao
Kempe, Julia
Munos, Remi
contents Reinforcement learning (RL) is increasingly used to align large language models (LLMs). Off-policy methods offer greater implementation simplicity and data efficiency than on-policy techniques, but often result in suboptimal performance. In this work, we study the intermediate range of algorithms between off-policy RL and supervised fine-tuning by analyzing a simple off-policy REINFORCE algorithm, where the advantage is defined as $A=r-V$, with $r$ a reward and $V$ some tunable baseline. Intuitively, lowering $V$ emphasizes high-reward samples, while raising it penalizes low-reward ones more heavily. We first provide a theoretical analysis of this off-policy REINFORCE algorithm, showing that when the baseline $V$ lower-bounds the expected reward, the algorithm enjoys a policy improvement guarantee. Our analysis reveals that while on-policy updates can safely leverage both positive and negative signals, off-policy updates benefit from focusing more on positive rewards than on negative ones. We validate our findings experimentally in a controlled stochastic bandit setting and through fine-tuning state-of-the-art LLMs on reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
Arnal, Charles
Narozniak, Gaëtan
Cabannes, Vivien
Tang, Yunhao
Kempe, Julia
Munos, Remi
Machine Learning
Computation and Language
Reinforcement learning (RL) is increasingly used to align large language models (LLMs). Off-policy methods offer greater implementation simplicity and data efficiency than on-policy techniques, but often result in suboptimal performance. In this work, we study the intermediate range of algorithms between off-policy RL and supervised fine-tuning by analyzing a simple off-policy REINFORCE algorithm, where the advantage is defined as $A=r-V$, with $r$ a reward and $V$ some tunable baseline. Intuitively, lowering $V$ emphasizes high-reward samples, while raising it penalizes low-reward ones more heavily. We first provide a theoretical analysis of this off-policy REINFORCE algorithm, showing that when the baseline $V$ lower-bounds the expected reward, the algorithm enjoys a policy improvement guarantee. Our analysis reveals that while on-policy updates can safely leverage both positive and negative signals, off-policy updates benefit from focusing more on positive rewards than on negative ones. We validate our findings experimentally in a controlled stochastic bandit setting and through fine-tuning state-of-the-art LLMs on reasoning tasks.
title Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.20520