GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912880178233344 |
|---|---|
| author | Tan, Hongze Wang, Zihan Pan, Jianfei Lin, Jinghao Wang, Hao Wu, Yifan Chen, Tao Zheng, Zhihang Tang, Zhihao Yang, Haihua |
| author_facet | Tan, Hongze Wang, Zihan Pan, Jianfei Lin, Jinghao Wang, Hao Wu, Yifan Chen, Tao Zheng, Zhihang Tang, Zhihao Yang, Haihua |
| contents | Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we propose Dynamic Entropy Weighting, systematically define entropy-based weight ratios $\frac{H_{i,t}}{\sum_{k=1}^{n} H_{k,t}}$ and similar variants to redistribute rewards and get fine-grained rewards through two new algorithms: Group Token Policy Optimization (GTPO), which assigns an entropy-weighted reward to each token and synthesizes token-specific advantage function to drive the model toward optimal path, and the analogous algorithm Sequence-Level GRPO (GRPO-S), which extends this design to the sequence level and exhibits superior stability in long Chain-of-Thought (CoT) reasoning tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_04349 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy Tan, Hongze Wang, Zihan Pan, Jianfei Lin, Jinghao Wang, Hao Wu, Yifan Chen, Tao Zheng, Zhihang Tang, Zhihao Yang, Haihua Computation and Language Artificial Intelligence Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we propose Dynamic Entropy Weighting, systematically define entropy-based weight ratios $\frac{H_{i,t}}{\sum_{k=1}^{n} H_{k,t}}$ and similar variants to redistribute rewards and get fine-grained rewards through two new algorithms: Group Token Policy Optimization (GTPO), which assigns an entropy-weighted reward to each token and synthesizes token-specific advantage function to drive the model toward optimal path, and the analogous algorithm Sequence-Level GRPO (GRPO-S), which extends this design to the sequence level and exhibits superior stability in long Chain-of-Thought (CoT) reasoning tasks. |
| title | GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2508.04349 |