VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Guobin, Zhao, Chenxiao, Cheng, Xiang, Huang, Lei, Yu, Xing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913100451545088
author Shen, Guobin
Zhao, Chenxiao
Cheng, Xiang
Huang, Lei
Yu, Xing
author_facet Shen, Guobin
Zhao, Chenxiao
Cheng, Xiang
Huang, Lei
Yu, Xing
contents Off-policy updates are inevitable in reinforcement learning (RL) for large language models (LLMs) due to rollout staleness from asynchronous training and mismatches between training and inference engines. Naive importance sampling gives an unbiased correction but suffers from high variance, which is amplified by unbounded ratios and autoregressive generation. Prior remedies either rely on scenario-specific engineering, or trade bias for variance via token-level clipping or sequence-level normalization, yet these approaches remain largely heuristic. We propose Variational sEquence-level Soft Policy Optimization (VESPO). By explicitly incorporating variance reduction into a variational formulation, we derive a principled closed-form reshaping kernel that operates directly on sequence-level importance weights, avoids token-level approximation and length normalization, and admits an explicit variance bound for the deployed kernel. Experiments on math reasoning and code generation show that VESPO maintains stable training under severe off-policy conditions (staleness up to 64x) and delivers consistent gains across both dense and Mixture-of-Experts (MoE) models, outperforming recent reshaping baselines under matched setup. Code is available at https://github.com/FloyedShen/VESPO.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10693
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Shen, Guobin
Zhao, Chenxiao
Cheng, Xiang
Huang, Lei
Yu, Xing
Machine Learning
Artificial Intelligence
Off-policy updates are inevitable in reinforcement learning (RL) for large language models (LLMs) due to rollout staleness from asynchronous training and mismatches between training and inference engines. Naive importance sampling gives an unbiased correction but suffers from high variance, which is amplified by unbounded ratios and autoregressive generation. Prior remedies either rely on scenario-specific engineering, or trade bias for variance via token-level clipping or sequence-level normalization, yet these approaches remain largely heuristic. We propose Variational sEquence-level Soft Policy Optimization (VESPO). By explicitly incorporating variance reduction into a variational formulation, we derive a principled closed-form reshaping kernel that operates directly on sequence-level importance weights, avoids token-level approximation and length normalization, and admits an explicit variance bound for the deployed kernel. Experiments on math reasoning and code generation show that VESPO maintains stable training under severe off-policy conditions (staleness up to 64x) and delivers consistent gains across both dense and Mixture-of-Experts (MoE) models, outperforming recent reshaping baselines under matched setup. Code is available at https://github.com/FloyedShen/VESPO.
title VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.10693