Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gong, Xue, Yi, Qi, Nan, Ziyuan, Huang, Guanhua, Li, Kejiao, Jiang, Yuhao, Xiong, Ruibin, Xu, Zenan, Guo, Jiaming, Peng, Shaohui, Zhou, Bo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917195984928768
author Gong, Xue
Yi, Qi
Nan, Ziyuan
Huang, Guanhua
Li, Kejiao
Jiang, Yuhao
Xiong, Ruibin
Xu, Zenan
Guo, Jiaming
Peng, Shaohui
Zhou, Bo
author_facet Gong, Xue
Yi, Qi
Nan, Ziyuan
Huang, Guanhua
Li, Kejiao
Jiang, Yuhao
Xiong, Ruibin
Xu, Zenan
Guo, Jiaming
Peng, Shaohui
Zhou, Bo
contents Training Large Language Models (LLMs) for reasoning tasks is increasingly driven by Reinforcement Learning with Verifiable Rewards (RLVR), where Proximal Policy Optimization (PPO) provides a principled framework for stable policy updates. However, the practical application of PPO is hindered by unreliable advantage estimation in the sparse-reward RLVR regime. This issue arises because the sparse rewards in RLVR lead to inaccurate intermediate value predictions, which in turn introduce significant bias when aggregated at every token by Generalized Advantage Estimation (GAE). To address this, we introduce Segmental Advantage Estimation (SAE), which mitigates the bias that GAE can incur in RLVR. Our key insight is that aggregating $n$-step advantages at every token(as in GAE) is unnecessary and often introduces excessive bias, since individual tokens carry minimal information. Instead, SAE first partitions the generated sequence into coherent sub-segments using low-probability tokens as heuristic boundaries. It then selectively computes variance-reduced advantage estimates only from these information-rich segment transitions, effectively filtering out noise from intermediate tokens. Our experiments demonstrate that SAE achieves superior performance, with marked improvements in final scores, training stability, and sample efficiency. These gains are shown to be consistent across multiple model sizes, and a correlation analysis confirms that our proposed advantage estimator achieves a higher correlation with an approximate ground-truth advantage, justifying its superior performance.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07320
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
Gong, Xue
Yi, Qi
Nan, Ziyuan
Huang, Guanhua
Li, Kejiao
Jiang, Yuhao
Xiong, Ruibin
Xu, Zenan
Guo, Jiaming
Peng, Shaohui
Zhou, Bo
Machine Learning
Artificial Intelligence
Training Large Language Models (LLMs) for reasoning tasks is increasingly driven by Reinforcement Learning with Verifiable Rewards (RLVR), where Proximal Policy Optimization (PPO) provides a principled framework for stable policy updates. However, the practical application of PPO is hindered by unreliable advantage estimation in the sparse-reward RLVR regime. This issue arises because the sparse rewards in RLVR lead to inaccurate intermediate value predictions, which in turn introduce significant bias when aggregated at every token by Generalized Advantage Estimation (GAE). To address this, we introduce Segmental Advantage Estimation (SAE), which mitigates the bias that GAE can incur in RLVR. Our key insight is that aggregating $n$-step advantages at every token(as in GAE) is unnecessary and often introduces excessive bias, since individual tokens carry minimal information. Instead, SAE first partitions the generated sequence into coherent sub-segments using low-probability tokens as heuristic boundaries. It then selectively computes variance-reduced advantage estimates only from these information-rich segment transitions, effectively filtering out noise from intermediate tokens. Our experiments demonstrate that SAE achieves superior performance, with marked improvements in final scores, training stability, and sample efficiency. These gains are shown to be consistent across multiple model sizes, and a correlation analysis confirms that our proposed advantage estimator achieves a higher correlation with an approximate ground-truth advantage, justifying its superior performance.
title Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.07320