EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Wujiang, Zhao, Wentian, Wang, Zhenting, Li, Yu-Jhe, Jin, Can, Jin, Mingyu, Mei, Kai, Wan, Kun, Metaxas, Dimitris N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914317203406848
author Xu, Wujiang
Zhao, Wentian
Wang, Zhenting
Li, Yu-Jhe
Jin, Can
Jin, Mingyu
Mei, Kai
Wan, Kun
Metaxas, Dimitris N.
author_facet Xu, Wujiang
Zhao, Wentian
Wang, Zhenting
Li, Yu-Jhe
Jin, Can
Jin, Mingyu
Mei, Kai
Wan, Kun
Metaxas, Dimitris N.
contents Training LLM agents in multi-turn environments with sparse rewards, where completing a single task requires 30+ turns of interaction within an episode, presents a fundamental challenge for reinforcement learning. We identify a critical failure mode unique to this setting: the exploration-exploitation cascade failure. This cascade begins with early-stage policy premature convergence, where sparse feedback causes agents to commit to flawed, low-entropy strategies. Subsequently, agents enter late-stage policy collapse, where conventional entropy regularization becomes counterproductive, promoting chaotic exploration that destabilizes training. We propose Entropy-regularized Policy Optimization (EPO), a general framework that breaks this failure cycle through three synergistic mechanisms: (1) adopting entropy regularization in multi-turn settings to enhance exploration, (2) an entropy smoothing regularizer that bounds policy entropy within historical averages to prevent abrupt fluctuations, and (3) adaptive phase-based weighting that balances exploration and exploitation across training. Our analysis justifies that EPO guarantees monotonically decreasing entropy variance while maintaining convergence. EPO achieves up to 152% performance improvement on ScienceWorld and up to 19.8% on ALFWorld. Our work demonstrates that multi-turn sparse-reward settings require fundamentally different entropy control than traditional RL, with broad implications for LLM agent training.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Xu, Wujiang
Zhao, Wentian
Wang, Zhenting
Li, Yu-Jhe
Jin, Can
Jin, Mingyu
Mei, Kai
Wan, Kun
Metaxas, Dimitris N.
Machine Learning
Computation and Language
Training LLM agents in multi-turn environments with sparse rewards, where completing a single task requires 30+ turns of interaction within an episode, presents a fundamental challenge for reinforcement learning. We identify a critical failure mode unique to this setting: the exploration-exploitation cascade failure. This cascade begins with early-stage policy premature convergence, where sparse feedback causes agents to commit to flawed, low-entropy strategies. Subsequently, agents enter late-stage policy collapse, where conventional entropy regularization becomes counterproductive, promoting chaotic exploration that destabilizes training. We propose Entropy-regularized Policy Optimization (EPO), a general framework that breaks this failure cycle through three synergistic mechanisms: (1) adopting entropy regularization in multi-turn settings to enhance exploration, (2) an entropy smoothing regularizer that bounds policy entropy within historical averages to prevent abrupt fluctuations, and (3) adaptive phase-based weighting that balances exploration and exploitation across training. Our analysis justifies that EPO guarantees monotonically decreasing entropy variance while maintaining convergence. EPO achieves up to 152% performance improvement on ScienceWorld and up to 19.8% on ALFWorld. Our work demonstrates that multi-turn sparse-reward settings require fundamentally different entropy control than traditional RL, with broad implications for LLM agent training.
title EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.22576