Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Yun, Wang, Qi, Mao, Yixiu, Zou, Heming, Jiang, Yuhang, Li, Yingyue, Xu, Wutong, Cai, Lizhou, Liu, Weijie, Bai, Clive, Yang, Kai, Chen, Yangkun, Yang, Saiyong, Ji, Xiangyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Doubly Mild Generalization for Offline Reinforcement Learning
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
by: Xu, Xin, et al.
Published: (2026)
by: Xu, Xin, et al.
Published: (2026)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
by: Mao, Yixiu, et al.
Published: (2025)
by: Mao, Yixiu, et al.
Published: (2025)
Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
by: Tang, Chenming, et al.
Published: (2025)
by: Tang, Chenming, et al.
Published: (2025)
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments
by: Qu, Yun, et al.
Published: (2025)
by: Qu, Yun, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Robust Fast Adaptation from Adversarially Explicit Task Distribution Generation
by: Wang, Cheems, et al.
Published: (2024)
by: Wang, Cheems, et al.
Published: (2024)
Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
by: Qu, Yun, et al.
Published: (2025)
by: Qu, Yun, et al.
Published: (2025)
Model Predictive Task Sampling for Efficient and Robust Adaptation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Structural features of the fly olfactory circuit mitigate the stability-plasticity dilemma in continual learning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Debiased Model-based Representations for Sample-efficient Continuous Control
by: Lyu, Jiafei, et al.
Published: (2026)
by: Lyu, Jiafei, et al.
Published: (2026)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
by: Liu, Fanfan, et al.
Published: (2026)
by: Liu, Fanfan, et al.
Published: (2026)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Nanotechnology‐Enabled Targeted Protein Degradation for Cancer Therapeutics
by: Wutong Zhao, et al.
Published: (2024)
by: Wutong Zhao, et al.
Published: (2024)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
by: Tang, Chenming, et al.
Published: (2026)
by: Tang, Chenming, et al.
Published: (2026)
Maestro: Learning to Collaborate via Conditional Listwise Policy Optimization for Multi-Agent LLMs
by: Yang, Wei, et al.
Published: (2025)
by: Yang, Wei, et al.
Published: (2025)
VAO: Validation-Aligned Optimization for Cross-Task Generative Auto-Bidding
by: Lv, Yiqin, et al.
Published: (2025)
by: Lv, Yiqin, et al.
Published: (2025)
LLM-Empowered State Representation for Reinforcement Learning
by: Wang, Boyuan, et al.
Published: (2024)
by: Wang, Boyuan, et al.
Published: (2024)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
On Extremal Volume Projections of the Simplex and the Cube
by: Pandis, Christos
Published: (2026)
by: Pandis, Christos
Published: (2026)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Post-Evental Epistemology: Consciousness, Temporality, and the Ontology of the Already-Ended
by: Sulin, Zhang, et al.
Published: (2025)
by: Sulin, Zhang, et al.
Published: (2025)
Enhancing Pretrained Model-based Continual Representation Learning via Guided Random Projection
by: Li, Ruilin, et al.
Published: (2026)
by: Li, Ruilin, et al.
Published: (2026)
Targeting C12ORF49 ‐Mediated Ferroptosis in Hepatocellular Carcinoma
by: Yuexin Liu, et al.
Published: (2026)
by: Yuexin Liu, et al.
Published: (2026)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Maiden Erlegh: An English Secondary School Development Project. Programme on Educational Building 2.
by: Booth, Clive
Published: (1973)
by: Booth, Clive
Published: (1973)
Similar Items
-
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026) -
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026) -
Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
by: Mao, Yixiu, et al.
Published: (2026) -
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025) -
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025)