GVPO: Group Variance Policy Optimization for Large Language Model Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaichen, Hong, Yuzhong, Bao, Junwei, Jiang, Hongfei, Song, Yang, Hong, Dingqian, Xiong, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917043843891200
author Zhang, Kaichen
Hong, Yuzhong
Bao, Junwei
Jiang, Hongfei
Song, Yang
Hong, Dingqian
Xiong, Hui
author_facet Zhang, Kaichen
Hong, Yuzhong
Bao, Junwei
Jiang, Hongfei
Song, Yang
Hong, Dingqian
Xiong, Hui
contents Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training techniques, such as Group Relative Policy Optimization (GRPO), leverage increased sampling with relative reward scoring to achieve superior performance, these methods often suffer from training instability that limits their practical adoption. As a next step, we present Group Variance Policy Optimization (GVPO). GVPO incorporates the analytical solution to KL-constrained reward maximization directly into its gradient weights, ensuring alignment with the optimal policy. The method provides intuitive physical interpretations: its gradient mirrors the mean squared error between the central distance of implicit rewards and that of actual rewards. GVPO offers two key advantages: (1) it guarantees a unique optimal solution, exactly the KL-constrained reward maximization objective, (2) it supports flexible sampling distributions that avoids on-policy and importance sampling limitations. By unifying theoretical guarantees with practical adaptability, GVPO establishes a new paradigm for reliable and versatile LLM post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
Zhang, Kaichen
Hong, Yuzhong
Bao, Junwei
Jiang, Hongfei
Song, Yang
Hong, Dingqian
Xiong, Hui
Artificial Intelligence
Machine Learning
Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training techniques, such as Group Relative Policy Optimization (GRPO), leverage increased sampling with relative reward scoring to achieve superior performance, these methods often suffer from training instability that limits their practical adoption. As a next step, we present Group Variance Policy Optimization (GVPO). GVPO incorporates the analytical solution to KL-constrained reward maximization directly into its gradient weights, ensuring alignment with the optimal policy. The method provides intuitive physical interpretations: its gradient mirrors the mean squared error between the central distance of implicit rewards and that of actual rewards. GVPO offers two key advantages: (1) it guarantees a unique optimal solution, exactly the KL-constrained reward maximization objective, (2) it supports flexible sampling distributions that avoids on-policy and importance sampling limitations. By unifying theoretical guarantees with practical adaptability, GVPO establishes a new paradigm for reliable and versatile LLM post-training.
title GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.19599