Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhihe, Luo, Xufang, Wang, Zilong, Han, Dongqi, He, Zhiyuan, Li, Dongsheng, Xu, Yunjian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915292617113600
author Yang, Zhihe
Luo, Xufang
Wang, Zilong
Han, Dongqi
He, Zhiyuan
Li, Dongsheng
Xu, Yunjian
author_facet Yang, Zhihe
Luo, Xufang
Wang, Zilong
Han, Dongqi
He, Zhiyuan
Li, Dongsheng
Xu, Yunjian
contents Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy Optimization (GRPO) demonstrating exceptional effectiveness. In this study, we identify a critical yet underexplored issue in RL training: low-probability tokens disproportionately influence model updates due to their large gradient magnitudes. This dominance hinders the effective learning of high-probability tokens, whose gradients are essential for LLMs' performance but are substantially suppressed. To mitigate this interference, we propose two novel methods: Advantage Reweighting and Low-Probability Token Isolation (Lopti), both of which effectively attenuate gradients from low-probability tokens while emphasizing parameter updates driven by high-probability tokens. Our approaches promote balanced updates across tokens with varying probabilities, thereby enhancing the efficiency of RL training. Experimental results demonstrate that they substantially improve the performance of GRPO-trained LLMs, achieving up to a 46.2% improvement in K&K Logic Puzzle reasoning tasks. Our implementation is available at https://github.com/zhyang2226/AR-Lopti.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12929
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Yang, Zhihe
Luo, Xufang
Wang, Zilong
Han, Dongqi
He, Zhiyuan
Li, Dongsheng
Xu, Yunjian
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy Optimization (GRPO) demonstrating exceptional effectiveness. In this study, we identify a critical yet underexplored issue in RL training: low-probability tokens disproportionately influence model updates due to their large gradient magnitudes. This dominance hinders the effective learning of high-probability tokens, whose gradients are essential for LLMs' performance but are substantially suppressed. To mitigate this interference, we propose two novel methods: Advantage Reweighting and Low-Probability Token Isolation (Lopti), both of which effectively attenuate gradients from low-probability tokens while emphasizing parameter updates driven by high-probability tokens. Our approaches promote balanced updates across tokens with varying probabilities, thereby enhancing the efficiency of RL training. Experimental results demonstrate that they substantially improve the performance of GRPO-trained LLMs, achieving up to a 46.2% improvement in K&K Logic Puzzle reasoning tasks. Our implementation is available at https://github.com/zhyang2226/AR-Lopti.
title Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.12929