SOUP: Token-level Single-sample Mix-policy Reinforcement Learning for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Lei, Bi, Wei, Sun, Chenxi, Jin, Renren, Xiong, Deyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910004483719168
author Yang, Lei
Bi, Wei
Sun, Chenxi
Jin, Renren
Xiong, Deyi
author_facet Yang, Lei
Bi, Wei
Sun, Chenxi
Jin, Renren
Xiong, Deyi
contents On-policy reinforcement learning (RL) methods widely used for language model post-training, like Group Relative Policy Optimization (GRPO), often suffer from limited exploration and early saturation due to low sampling diversity. While off-policy data can help, current approaches that mix entire trajectories cause significant policy mismatch and instability. In this work, we propose the $\textbf{S}$ingle-sample Mix-p$\textbf{O}$licy $\textbf{U}$nified $\textbf{P}$aradigm (SOUP), a framework that unifies off- and on-policy learning within individual samples at the token level. It confines off-policy influence to the prefix of a generated sequence sampled from historical policies, while the continuation is generated on-policy. Through token-level importance ratios, SOUP effectively leverages off-policy information while preserving training stability. Extensive experiments demonstrate that SOUP consistently outperforms standard on-policy training and existing off-policy extensions. Our further analysis clarifies how our fine-grained, single-sample mix-policy training can improve both exploration and final performance in LLM RL.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21476
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SOUP: Token-level Single-sample Mix-policy Reinforcement Learning for Large Language Models
Yang, Lei
Bi, Wei
Sun, Chenxi
Jin, Renren
Xiong, Deyi
Computation and Language
On-policy reinforcement learning (RL) methods widely used for language model post-training, like Group Relative Policy Optimization (GRPO), often suffer from limited exploration and early saturation due to low sampling diversity. While off-policy data can help, current approaches that mix entire trajectories cause significant policy mismatch and instability. In this work, we propose the $\textbf{S}$ingle-sample Mix-p$\textbf{O}$licy $\textbf{U}$nified $\textbf{P}$aradigm (SOUP), a framework that unifies off- and on-policy learning within individual samples at the token level. It confines off-policy influence to the prefix of a generated sequence sampled from historical policies, while the continuation is generated on-policy. Through token-level importance ratios, SOUP effectively leverages off-policy information while preserving training stability. Extensive experiments demonstrate that SOUP consistently outperforms standard on-policy training and existing off-policy extensions. Our further analysis clarifies how our fine-grained, single-sample mix-policy training can improve both exploration and final performance in LLM RL.
title SOUP: Token-level Single-sample Mix-policy Reinforcement Learning for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.21476