Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hao, Zhezheng, Wang, Hong, Liu, Haoyang, Luo, Jian, Yu, Jiarui, Dong, Hande, Lin, Qiang, Wang, Can, Chen, Jiawei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917446259048448
author Hao, Zhezheng
Wang, Hong
Liu, Haoyang
Luo, Jian
Yu, Jiarui
Dong, Hande
Lin, Qiang
Wang, Can
Chen, Jiawei
author_facet Hao, Zhezheng
Wang, Hong
Liu, Haoyang
Luo, Jian
Yu, Jiarui
Dong, Hande
Lin, Qiang
Wang, Can
Chen, Jiawei
contents Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, its training is often plagued by \emph{entropy collapse}, a rapid decline in policy entropy that limits exploration and undermines training effectiveness. While recent works attempt to mitigate this issue via several heuristic entropy interventions, the underlying mechanisms remain poorly understood. In this work, we conduct comprehensive theoretical and empirical analyses of entropy dynamics in RLVR, offering two main insights: (1) We derive a tight analytical approximation for token-level entropy change at each update step, revealing four governing factors and providing a unified theoretical framework to explain how existing methods influence entropy; (2) We reveal a fundamental limitation of recent approaches: they rely on heuristic adjustments to one or two of these factors, leaving other relevant factors unconsidered, thus inherently limiting their effectiveness. Motivated by these findings, we propose STEER, a principled entropy-modulation method that adaptively reweights tokens based on theoretically-estimated entropy variations. Extensive experiments across six mathematical reasoning and three coding benchmarks demonstrate that STEER effectively mitigates entropy collapse and consistently outperforms state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10150
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Hao, Zhezheng
Wang, Hong
Liu, Haoyang
Luo, Jian
Yu, Jiarui
Dong, Hande
Lin, Qiang
Wang, Can
Chen, Jiawei
Machine Learning
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, its training is often plagued by \emph{entropy collapse}, a rapid decline in policy entropy that limits exploration and undermines training effectiveness. While recent works attempt to mitigate this issue via several heuristic entropy interventions, the underlying mechanisms remain poorly understood. In this work, we conduct comprehensive theoretical and empirical analyses of entropy dynamics in RLVR, offering two main insights: (1) We derive a tight analytical approximation for token-level entropy change at each update step, revealing four governing factors and providing a unified theoretical framework to explain how existing methods influence entropy; (2) We reveal a fundamental limitation of recent approaches: they rely on heuristic adjustments to one or two of these factors, leaving other relevant factors unconsidered, thus inherently limiting their effectiveness. Motivated by these findings, we propose STEER, a principled entropy-modulation method that adaptively reweights tokens based on theoretically-estimated entropy variations. Extensive experiments across six mathematical reasoning and three coding benchmarks demonstrate that STEER effectively mitigates entropy collapse and consistently outperforms state-of-the-art baselines.
title Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.10150