Reasoning with Exploration: An Entropy Perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Daixuan, Huang, Shaohan, Zhu, Xuekai, Dai, Bo, Zhao, Wayne Xin, Zhang, Zhenliang, Wei, Furu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908637165780992
author Cheng, Daixuan
Huang, Shaohan
Zhu, Xuekai
Dai, Bo
Zhao, Wayne Xin
Zhang, Zhenliang
Wei, Furu
author_facet Cheng, Daixuan
Huang, Shaohan
Zhu, Xuekai
Dai, Bo
Zhao, Wayne Xin
Zhang, Zhenliang
Wei, Furu
contents Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing large language model (LLM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LLMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LLMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting longer and deeper reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LLM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LLM reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning with Exploration: An Entropy Perspective
Cheng, Daixuan
Huang, Shaohan
Zhu, Xuekai
Dai, Bo
Zhao, Wayne Xin
Zhang, Zhenliang
Wei, Furu
Computation and Language
Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing large language model (LLM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LLMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LLMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting longer and deeper reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LLM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LLM reasoning.
title Reasoning with Exploration: An Entropy Perspective
topic Computation and Language
url https://arxiv.org/abs/2506.14758