The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Agarwal, Shivam, Zhang, Zimin, Yuan, Lifan, Han, Jiawei, Peng, Hao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915295753404416
author Agarwal, Shivam
Zhang, Zimin
Yuan, Lifan
Han, Jiawei
Peng, Hao
author_facet Agarwal, Shivam
Zhang, Zimin
Yuan, Lifan
Han, Jiawei
Peng, Hao
contents Entropy minimization (EM) trains the model to concentrate even more probability mass on its most confident outputs. We show that this simple objective alone, without any labeled data, can substantially improve large language models' (LLMs) performance on challenging math, physics, and coding tasks. We explore three approaches: (1) EM-FT minimizes token-level entropy similarly to instruction finetuning, but on unlabeled outputs drawn from the model; (2) EM-RL: reinforcement learning with negative entropy as the only reward to maximize; (3) EM-INF: inference-time logit adjustment to reduce entropy without any training data or parameter updates. On Qwen-7B, EM-RL, without any labeled data, achieves comparable or better performance than strong RL baselines such as GRPO and RLOO that are trained on 60K labeled examples. Furthermore, EM-INF enables Qwen-32B to match or exceed the performance of proprietary models like GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro on the challenging SciCode benchmark, while being 3x more efficient than self-consistency and sequential refinement. Our findings reveal that many pretrained LLMs possess previously underappreciated reasoning capabilities that can be effectively elicited through entropy minimization alone, without any labeled data or even any parameter updates.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15134
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Agarwal, Shivam
Zhang, Zimin
Yuan, Lifan
Han, Jiawei
Peng, Hao
Machine Learning
Artificial Intelligence
Entropy minimization (EM) trains the model to concentrate even more probability mass on its most confident outputs. We show that this simple objective alone, without any labeled data, can substantially improve large language models' (LLMs) performance on challenging math, physics, and coding tasks. We explore three approaches: (1) EM-FT minimizes token-level entropy similarly to instruction finetuning, but on unlabeled outputs drawn from the model; (2) EM-RL: reinforcement learning with negative entropy as the only reward to maximize; (3) EM-INF: inference-time logit adjustment to reduce entropy without any training data or parameter updates. On Qwen-7B, EM-RL, without any labeled data, achieves comparable or better performance than strong RL baselines such as GRPO and RLOO that are trained on 60K labeled examples. Furthermore, EM-INF enables Qwen-32B to match or exceed the performance of proprietary models like GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro on the challenging SciCode benchmark, while being 3x more efficient than self-consistency and sequential refinement. Our findings reveal that many pretrained LLMs possess previously underappreciated reasoning capabilities that can be effectively elicited through entropy minimization alone, without any labeled data or even any parameter updates.
title The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.15134