Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Peter, Li, Xiaopeng, Li, Ziniu, Chen, Xi, Lin, Tianyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911502396555264
author Chen, Peter
Li, Xiaopeng
Li, Ziniu
Chen, Xi
Lin, Tianyi
author_facet Chen, Peter
Li, Xiaopeng
Li, Ziniu
Chen, Xi
Lin, Tianyi
contents Reinforcement learning (RL) has proven effective in strengthening the reasoning capabilities of large language models (LLMs). A widely adopted method, Group Relative Policy Optimization (GRPO), has shown strong empirical results in training recent reasoning models, but it fails to update the policy when all responses within a group are incorrect (i.e., all-negative-sample groups). This limitation highlights a gap between artificial and human intelligence: unlike humans, who can learn from mistakes, GRPO discards these failure signals. We introduce a simple framework to mitigate the all-negative-sample issue by incorporating response diversity within groups using a step-wise judge model, which can be trained directly or adapted from existing LLMs. In a simplified setting, we prove that this diversification accelerates GRPO's learning dynamics. We then empirically validate Stepwise Guided Policy Optimization (SGPO) across model sizes (7B, 14B, 32B) in both offline and online training on nine reasoning benchmarks (including base and distilled variants). Overall, SGPO improves average performance and is effective in early and mid-training when all-negative groups are prevalent, while improvements are not uniform across every benchmark and depend on the structure and informativeness of negative samples. Finally, SGPO does not require the judge model to generate correct solutions, distinguishing it from knowledge distillation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11595
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
Chen, Peter
Li, Xiaopeng
Li, Ziniu
Chen, Xi
Lin, Tianyi
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement learning (RL) has proven effective in strengthening the reasoning capabilities of large language models (LLMs). A widely adopted method, Group Relative Policy Optimization (GRPO), has shown strong empirical results in training recent reasoning models, but it fails to update the policy when all responses within a group are incorrect (i.e., all-negative-sample groups). This limitation highlights a gap between artificial and human intelligence: unlike humans, who can learn from mistakes, GRPO discards these failure signals. We introduce a simple framework to mitigate the all-negative-sample issue by incorporating response diversity within groups using a step-wise judge model, which can be trained directly or adapted from existing LLMs. In a simplified setting, we prove that this diversification accelerates GRPO's learning dynamics. We then empirically validate Stepwise Guided Policy Optimization (SGPO) across model sizes (7B, 14B, 32B) in both offline and online training on nine reasoning benchmarks (including base and distilled variants). Overall, SGPO improves average performance and is effective in early and mid-training when all-negative groups are prevalent, while improvements are not uniform across every benchmark and depend on the structure and informativeness of negative samples. Finally, SGPO does not require the judge model to generate correct solutions, distinguishing it from knowledge distillation methods.
title Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.11595