GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Hongze, Wang, Zihan, Pan, Jianfei, Lin, Jinghao, Wang, Hao, Wu, Yifan, Chen, Tao, Zheng, Zhihang, Tang, Zhihao, Yang, Haihua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912880178233344
author Tan, Hongze
Wang, Zihan
Pan, Jianfei
Lin, Jinghao
Wang, Hao
Wu, Yifan
Chen, Tao
Zheng, Zhihang
Tang, Zhihao
Yang, Haihua
author_facet Tan, Hongze
Wang, Zihan
Pan, Jianfei
Lin, Jinghao
Wang, Hao
Wu, Yifan
Chen, Tao
Zheng, Zhihang
Tang, Zhihao
Yang, Haihua
contents Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we propose Dynamic Entropy Weighting, systematically define entropy-based weight ratios $\frac{H_{i,t}}{\sum_{k=1}^{n} H_{k,t}}$ and similar variants to redistribute rewards and get fine-grained rewards through two new algorithms: Group Token Policy Optimization (GTPO), which assigns an entropy-weighted reward to each token and synthesizes token-specific advantage function to drive the model toward optimal path, and the analogous algorithm Sequence-Level GRPO (GRPO-S), which extends this design to the sequence level and exhibits superior stability in long Chain-of-Thought (CoT) reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Tan, Hongze
Wang, Zihan
Pan, Jianfei
Lin, Jinghao
Wang, Hao
Wu, Yifan
Chen, Tao
Zheng, Zhihang
Tang, Zhihao
Yang, Haihua
Computation and Language
Artificial Intelligence
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we propose Dynamic Entropy Weighting, systematically define entropy-based weight ratios $\frac{H_{i,t}}{\sum_{k=1}^{n} H_{k,t}}$ and similar variants to redistribute rewards and get fine-grained rewards through two new algorithms: Group Token Policy Optimization (GTPO), which assigns an entropy-weighted reward to each token and synthesizes token-specific advantage function to drive the model toward optimal path, and the analogous algorithm Sequence-Level GRPO (GRPO-S), which extends this design to the sequence level and exhibits superior stability in long Chain-of-Thought (CoT) reasoning tasks.
title GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.04349