Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jin, Renren, Gao, Pengzhi, Ren, Yuqi, Han, Zhuowen, Zhang, Tongxuan, Huang, Wuwei, Liu, Wei, Luan, Jian, Xiong, Deyi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910145009680384
author Jin, Renren
Gao, Pengzhi
Ren, Yuqi
Han, Zhuowen
Zhang, Tongxuan
Huang, Wuwei
Liu, Wei
Luan, Jian
Xiong, Deyi
author_facet Jin, Renren
Gao, Pengzhi
Ren, Yuqi
Han, Zhuowen
Zhang, Tongxuan
Huang, Wuwei
Liu, Wei
Luan, Jian
Xiong, Deyi
contents Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, the entropy of LLMs usually collapses during RLVR training, leading to premature convergence to suboptimal local minima and hindering further performance improvement. Although various approaches have been proposed to mitigate entropy collapse, a comprehensive study of entropy in RLVR remains lacking. To bridge this gap, we conduct extensive experiments to investigate the entropy dynamics of LLMs trained with RLVR and analyze how model entropy correlates with response diversity, calibration, and performance across various benchmarks. Our results identify three key factors that influence entropy: the clipping thresholds in the optimization objective, the number of off-policy updates, and the diversity of the training data. Furthermore, through both theoretical analysis and empirical validation, we demonstrate that tokens with positive advantages are the primary drivers of entropy collapse. Motivated by this insight, we propose Positive-Advantage Reweighting, a simple yet effective approach that regulates model entropy by adjusting the loss weights assigned to tokens with positive advantages during RLVR training, while maintaining competitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05993
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
Jin, Renren
Gao, Pengzhi
Ren, Yuqi
Han, Zhuowen
Zhang, Tongxuan
Huang, Wuwei
Liu, Wei
Luan, Jian
Xiong, Deyi
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, the entropy of LLMs usually collapses during RLVR training, leading to premature convergence to suboptimal local minima and hindering further performance improvement. Although various approaches have been proposed to mitigate entropy collapse, a comprehensive study of entropy in RLVR remains lacking. To bridge this gap, we conduct extensive experiments to investigate the entropy dynamics of LLMs trained with RLVR and analyze how model entropy correlates with response diversity, calibration, and performance across various benchmarks. Our results identify three key factors that influence entropy: the clipping thresholds in the optimization objective, the number of off-policy updates, and the diversity of the training data. Furthermore, through both theoretical analysis and empirical validation, we demonstrate that tokens with positive advantages are the primary drivers of entropy collapse. Motivated by this insight, we propose Positive-Advantage Reweighting, a simple yet effective approach that regulates model entropy by adjusting the loss weights assigned to tokens with positive advantages during RLVR training, while maintaining competitive performance.
title Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.05993