Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Yun, Wang, Qi, Mao, Yixiu, Zou, Heming, Jiang, Yuhang, Li, Yingyue, Xu, Wutong, Cai, Lizhou, Liu, Weijie, Bai, Clive, Yang, Kai, Chen, Yangkun, Yang, Saiyong, Ji, Xiangyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911699600146432
author Qu, Yun
Wang, Qi
Mao, Yixiu
Zou, Heming
Jiang, Yuhang
Li, Yingyue
Xu, Wutong
Cai, Lizhou
Liu, Weijie
Bai, Clive
Yang, Kai
Chen, Yangkun
Yang, Saiyong
Ji, Xiangyang
author_facet Qu, Yun
Wang, Qi
Mao, Yixiu
Zou, Heming
Jiang, Yuhang
Li, Yingyue
Xu, Wutong
Cai, Lizhou
Liu, Weijie
Bai, Clive
Yang, Kai
Chen, Yangkun
Yang, Saiyong
Ji, Xiangyang
contents Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning capacity. Among existing recipes, group-based policy gradient is prevalent, which samples a group of responses per prompt and updates the policy via group-relative advantage signals. This work reveals that these optimization strategies share a common geometric structure: each implicitly defines a target distribution on the response simplex and projects toward it via first-order approximation. Building on this insight, we propose Listwise Policy Optimization (LPO) to explicitly conduct the target-projection, which demystifies the implicit target by restricting the proximal RL objective to the response simplex, and then projects the policy via exact divergence minimization. This framework provides (i) monotonic improvement on the listwise objective with bounded, zero-sum, and self-correcting projection gradients, and (ii) flexibility in divergence selection with distinct structural properties through the decoupled projection step. On diverse reasoning tasks and LLM backbones, LPO consistently improves training performance over typical policy gradient baselines under matched targets, while intrinsically preserving optimization stability and response diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
Qu, Yun
Wang, Qi
Mao, Yixiu
Zou, Heming
Jiang, Yuhang
Li, Yingyue
Xu, Wutong
Cai, Lizhou
Liu, Weijie
Bai, Clive
Yang, Kai
Chen, Yangkun
Yang, Saiyong
Ji, Xiangyang
Machine Learning
Artificial Intelligence
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning capacity. Among existing recipes, group-based policy gradient is prevalent, which samples a group of responses per prompt and updates the policy via group-relative advantage signals. This work reveals that these optimization strategies share a common geometric structure: each implicitly defines a target distribution on the response simplex and projects toward it via first-order approximation. Building on this insight, we propose Listwise Policy Optimization (LPO) to explicitly conduct the target-projection, which demystifies the implicit target by restricting the proximal RL objective to the response simplex, and then projects the policy via exact divergence minimization. This framework provides (i) monotonic improvement on the listwise objective with bounded, zero-sum, and self-correcting projection gradients, and (ii) flexibility in divergence selection with distinct structural properties through the decoupled projection step. On diverse reasoning tasks and LLM backbones, LPO consistently improves training performance over typical policy gradient baselines under matched targets, while intrinsically preserving optimization stability and response diversity.
title Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.06139