Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chen, Liu, Nazhou, Yang, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814109503488
author Li, Chen
Liu, Nazhou
Yang, Kai
author_facet Li, Chen
Liu, Nazhou
Yang, Kai
contents Since DeepSeek-R1 popularized, Group Relative Policy Optimization (GRPO) has become the core part of training Reasoning LLMs. However, we find some deficiency that influences RL stability and inference efficiency, like zero-variance in advantage estimation. Thus, we propose Adaptive Group Policy Optimization (AGPO) which uses a simple but effective method, an adaptive loss function, to mitigate training fluctuation and token inefficiency. The experiments demonstrate our method achieves more stable training and superior performance with significantly fewer tokens in reasoning steps.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
Li, Chen
Liu, Nazhou
Yang, Kai
Computation and Language
Since DeepSeek-R1 popularized, Group Relative Policy Optimization (GRPO) has become the core part of training Reasoning LLMs. However, we find some deficiency that influences RL stability and inference efficiency, like zero-variance in advantage estimation. Thus, we propose Adaptive Group Policy Optimization (AGPO) which uses a simple but effective method, an adaptive loss function, to mitigate training fluctuation and token inefficiency. The experiments demonstrate our method achieves more stable training and superior performance with significantly fewer tokens in reasoning steps.
title Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
topic Computation and Language
url https://arxiv.org/abs/2503.15952