Group Sequence Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Chujie, Liu, Shixuan, Li, Mingze, Chen, Xiong-Hui, Yu, Bowen, Gao, Chang, Dang, Kai, Liu, Yuqiong, Men, Rui, Yang, An, Zhou, Jingren, Lin, Junyang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915413147779072
author Zheng, Chujie
Liu, Shixuan
Li, Mingze
Chen, Xiong-Hui
Yu, Bowen
Gao, Chang
Dang, Kai
Liu, Yuqiong
Men, Rui
Yang, An
Zhou, Jingren
Lin, Junyang
author_facet Zheng, Chujie
Liu, Shixuan
Li, Mingze
Chen, Xiong-Hui
Yu, Bowen
Gao, Chang
Dang, Kai
Liu, Yuqiong
Men, Rui
Yang, An
Zhou, Jingren
Lin, Junyang
contents This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models. Unlike previous algorithms that adopt token-level importance ratios, GSPO defines the importance ratio based on sequence likelihood and performs sequence-level clipping, rewarding, and optimization. We demonstrate that GSPO achieves superior training efficiency and performance compared to the GRPO algorithm, notably stabilizes Mixture-of-Experts (MoE) RL training, and has the potential for simplifying the design of RL infrastructure. These merits of GSPO have contributed to the remarkable improvements in the latest Qwen3 models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group Sequence Policy Optimization
Zheng, Chujie
Liu, Shixuan
Li, Mingze
Chen, Xiong-Hui
Yu, Bowen
Gao, Chang
Dang, Kai
Liu, Yuqiong
Men, Rui
Yang, An
Zhou, Jingren
Lin, Junyang
Machine Learning
Artificial Intelligence
Computation and Language
This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models. Unlike previous algorithms that adopt token-level importance ratios, GSPO defines the importance ratio based on sequence likelihood and performs sequence-level clipping, rewarding, and optimization. We demonstrate that GSPO achieves superior training efficiency and performance compared to the GRPO algorithm, notably stabilizes Mixture-of-Experts (MoE) RL training, and has the potential for simplifying the design of RL infrastructure. These merits of GSPO have contributed to the remarkable improvements in the latest Qwen3 models.
title Group Sequence Policy Optimization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.18071