Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zikang, Yue, Tongtian, Tang, Yepeng, Guo, Longteng, Cai, Junxian, Liu, Qingbin, Chen, Xi, Liu, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908395419729920
author Liu, Zikang
Yue, Tongtian
Tang, Yepeng
Guo, Longteng
Cai, Junxian
Liu, Qingbin
Chen, Xi
Liu, Jing
author_facet Liu, Zikang
Yue, Tongtian
Tang, Yepeng
Guo, Longteng
Cai, Junxian
Liu, Qingbin
Chen, Xi
Liu, Jing
contents Group Relative Policy Optimization (GRPO) enhances policy learning by computing gradients from relative comparisons among candidate outputs that share a common input prefix. Despite its effectiveness, GRPO introduces substantial computational overhead when processing long shared prefixes, which must be redundantly encoded for each group member. This inefficiency becomes a major scalability bottleneck in long-context learning scenarios. We propose Prefix Grouper, an efficient GRPO training algorithm that eliminates redundant prefix computation via a Shared-Prefix Forward strategy. In particular, by restructuring self-attention into two parts, our method enables the shared prefix to be encoded only once, while preserving full differentiability and compatibility with end-to-end training. We provide both theoretical and empirical evidence that Prefix Grouper is training-equivalent to standard GRPO: it yields identical forward outputs and backward gradients, ensuring that the optimization dynamics and final policy performance remain unchanged. Empirically, our experiments confirm that Prefix Grouper achieves consistent results while significantly reducing the computational cost of training, particularly in long-prefix scenarios. The proposed method is fully plug-and-play: it is compatible with existing GRPO-based architectures and can be seamlessly integrated into current training pipelines as a drop-in replacement, requiring no structural modifications and only minimal changes to input construction and attention computation. Prefix Grouper enables the use of larger group sizes under the same computational budget, thereby improving the scalability of GRPO to more complex tasks and larger models. Code is now available at https://github.com/johncaged/PrefixGrouper
format Preprint
id arxiv_https___arxiv_org_abs_2506_05433
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
Liu, Zikang
Yue, Tongtian
Tang, Yepeng
Guo, Longteng
Cai, Junxian
Liu, Qingbin
Chen, Xi
Liu, Jing
Machine Learning
Artificial Intelligence
Group Relative Policy Optimization (GRPO) enhances policy learning by computing gradients from relative comparisons among candidate outputs that share a common input prefix. Despite its effectiveness, GRPO introduces substantial computational overhead when processing long shared prefixes, which must be redundantly encoded for each group member. This inefficiency becomes a major scalability bottleneck in long-context learning scenarios. We propose Prefix Grouper, an efficient GRPO training algorithm that eliminates redundant prefix computation via a Shared-Prefix Forward strategy. In particular, by restructuring self-attention into two parts, our method enables the shared prefix to be encoded only once, while preserving full differentiability and compatibility with end-to-end training. We provide both theoretical and empirical evidence that Prefix Grouper is training-equivalent to standard GRPO: it yields identical forward outputs and backward gradients, ensuring that the optimization dynamics and final policy performance remain unchanged. Empirically, our experiments confirm that Prefix Grouper achieves consistent results while significantly reducing the computational cost of training, particularly in long-prefix scenarios. The proposed method is fully plug-and-play: it is compatible with existing GRPO-based architectures and can be seamlessly integrated into current training pipelines as a drop-in replacement, requiring no structural modifications and only minimal changes to input construction and attention computation. Prefix Grouper enables the use of larger group sizes under the same computational budget, thereby improving the scalability of GRPO to more complex tasks and larger models. Code is now available at https://github.com/johncaged/PrefixGrouper
title Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.05433