The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cui, Ganqu, Zhang, Yuchen, Chen, Jiacheng, Yuan, Lifan, Wang, Zhi, Zuo, Yuxin, Li, Haozhan, Fan, Yuchen, Chen, Huayu, Chen, Weize, Liu, Zhiyuan, Peng, Hao, Bai, Lei, Ouyang, Wanli, Cheng, Yu, Zhou, Bowen, Ding, Ning
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908382962647040
author Cui, Ganqu
Zhang, Yuchen
Chen, Jiacheng
Yuan, Lifan
Wang, Zhi
Zuo, Yuxin
Li, Haozhan
Fan, Yuchen
Chen, Huayu
Chen, Weize
Liu, Zhiyuan
Peng, Hao
Bai, Lei
Ouyang, Wanli
Cheng, Yu
Zhou, Bowen
Ding, Ning
author_facet Cui, Ganqu
Zhang, Yuchen
Chen, Jiacheng
Yuan, Lifan
Wang, Zhi
Zuo, Yuxin
Li, Haozhan
Fan, Yuchen
Chen, Huayu
Chen, Weize
Liu, Zhiyuan
Peng, Hao
Bai, Lei
Ouyang, Wanli
Cheng, Yu
Zhou, Bowen
Ding, Ning
contents This paper aims to overcome a major obstacle in scaling RL for reasoning with LLMs, namely the collapse of policy entropy. Such phenomenon is consistently observed across vast RL runs without entropy intervention, where the policy entropy dropped sharply at the early training stage, this diminished exploratory ability is always accompanied with the saturation of policy performance. In practice, we establish a transformation equation R=-a*e^H+b between entropy H and downstream performance R. This empirical law strongly indicates that, the policy performance is traded from policy entropy, thus bottlenecked by its exhaustion, and the ceiling is fully predictable H=0, R=-a+b. Our finding necessitates entropy management for continuous exploration toward scaling compute for RL. To this end, we investigate entropy dynamics both theoretically and empirically. Our derivation highlights that, the change in policy entropy is driven by the covariance between action probability and the change in logits, which is proportional to its advantage when using Policy Gradient-like algorithms. Empirical study shows that, the values of covariance term and entropy differences matched exactly, supporting the theoretical conclusion. Moreover, the covariance term stays mostly positive throughout training, further explaining why policy entropy would decrease monotonically. Through understanding the mechanism behind entropy dynamics, we motivate to control entropy by restricting the update of high-covariance tokens. Specifically, we propose two simple yet effective techniques, namely Clip-Cov and KL-Cov, which clip and apply KL penalty to tokens with high covariances respectively. Experiments show that these methods encourage exploration, thus helping policy escape entropy collapse and achieve better downstream performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22617
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Cui, Ganqu
Zhang, Yuchen
Chen, Jiacheng
Yuan, Lifan
Wang, Zhi
Zuo, Yuxin
Li, Haozhan
Fan, Yuchen
Chen, Huayu
Chen, Weize
Liu, Zhiyuan
Peng, Hao
Bai, Lei
Ouyang, Wanli
Cheng, Yu
Zhou, Bowen
Ding, Ning
Machine Learning
Artificial Intelligence
Computation and Language
This paper aims to overcome a major obstacle in scaling RL for reasoning with LLMs, namely the collapse of policy entropy. Such phenomenon is consistently observed across vast RL runs without entropy intervention, where the policy entropy dropped sharply at the early training stage, this diminished exploratory ability is always accompanied with the saturation of policy performance. In practice, we establish a transformation equation R=-a*e^H+b between entropy H and downstream performance R. This empirical law strongly indicates that, the policy performance is traded from policy entropy, thus bottlenecked by its exhaustion, and the ceiling is fully predictable H=0, R=-a+b. Our finding necessitates entropy management for continuous exploration toward scaling compute for RL. To this end, we investigate entropy dynamics both theoretically and empirically. Our derivation highlights that, the change in policy entropy is driven by the covariance between action probability and the change in logits, which is proportional to its advantage when using Policy Gradient-like algorithms. Empirical study shows that, the values of covariance term and entropy differences matched exactly, supporting the theoretical conclusion. Moreover, the covariance term stays mostly positive throughout training, further explaining why policy entropy would decrease monotonically. Through understanding the mechanism behind entropy dynamics, we motivate to control entropy by restricting the update of high-covariance tokens. Specifically, we propose two simple yet effective techniques, namely Clip-Cov and KL-Cov, which clip and apply KL penalty to tokens with high covariances respectively. Experiments show that these methods encourage exploration, thus helping policy escape entropy collapse and achieve better downstream performance.
title The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.22617