Rethinking Entropy Regularization in Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yuxian, Li, Yafu, Chen, Guanxu, Liu, Dongrui, Cheng, Yu, Shao, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183899983872
author Jiang, Yuxian
Li, Yafu
Chen, Guanxu
Liu, Dongrui
Cheng, Yu
Shao, Jing
author_facet Jiang, Yuxian
Li, Yafu
Chen, Guanxu
Liu, Dongrui
Cheng, Yu
Shao, Jing
contents Reinforcement learning with verifiable rewards (RLVR) has shown great promise in enhancing the reasoning abilities of large reasoning models (LRMs). However, it suffers from a critical issue: entropy collapse and premature convergence. Naive entropy regularization, a common approach for encouraging exploration in the traditional RL literature, fails to address this problem in the context of LRM. Our analysis reveals that this failure stems from the vast action space and long trajectories in LRMs, which easily trigger a global entropy explosion as the model indiscriminately explores all possible actions and states. To address this, we propose SIREN (SelectIve entRopy rEgularizatioN), a method that confines exploration to a meaningful subset of actions and states. SIREN achieves this through a two-step entropy masking mechanism, consisting of a top-p mask and a peak-entropy mask. In addition, regularization is transformed into a self-anchored form to stabilize training. Across five mathematical benchmarks, SIREN attains superior average performance over previous entropy-related RLVR approaches, exemplified by a +6.6 maj@k improvement on AIME24/25 with Qwen2.5-Math-7B. Further analysis confirms that SIREN promotes greater response diversity and maintains entropy at an appropriate level, which helps to preserve the validation pass@k throughout training. This effectively mitigates the premature convergence problem common in RLVR for LRM.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Entropy Regularization in Large Reasoning Models
Jiang, Yuxian
Li, Yafu
Chen, Guanxu
Liu, Dongrui
Cheng, Yu
Shao, Jing
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement learning with verifiable rewards (RLVR) has shown great promise in enhancing the reasoning abilities of large reasoning models (LRMs). However, it suffers from a critical issue: entropy collapse and premature convergence. Naive entropy regularization, a common approach for encouraging exploration in the traditional RL literature, fails to address this problem in the context of LRM. Our analysis reveals that this failure stems from the vast action space and long trajectories in LRMs, which easily trigger a global entropy explosion as the model indiscriminately explores all possible actions and states. To address this, we propose SIREN (SelectIve entRopy rEgularizatioN), a method that confines exploration to a meaningful subset of actions and states. SIREN achieves this through a two-step entropy masking mechanism, consisting of a top-p mask and a peak-entropy mask. In addition, regularization is transformed into a self-anchored form to stabilize training. Across five mathematical benchmarks, SIREN attains superior average performance over previous entropy-related RLVR approaches, exemplified by a +6.6 maj@k improvement on AIME24/25 with Qwen2.5-Math-7B. Further analysis confirms that SIREN promotes greater response diversity and maintains entropy at an appropriate level, which helps to preserve the validation pass@k throughout training. This effectively mitigates the premature convergence problem common in RLVR for LRM.
title Rethinking Entropy Regularization in Large Reasoning Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.25133