Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Ruipeng, Chang, Ya-Chien, Gao, Sicun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2506.05615
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908404757299200
author Zhang, Ruipeng
Chang, Ya-Chien
Gao, Sicun
author_facet Zhang, Ruipeng
Chang, Ya-Chien
Gao, Sicun
contents The Maximum Entropy Reinforcement Learning (MaxEnt RL) framework is a leading approach for achieving efficient learning and robust performance across many RL tasks. However, MaxEnt methods have also been shown to struggle with performance-critical control problems in practice, where non-MaxEnt algorithms can successfully learn. In this work, we analyze how the trade-off between robustness and optimality affects the performance of MaxEnt algorithms in complex control tasks: while entropy maximization enhances exploration and robustness, it can also mislead policy optimization, leading to failure in tasks that require precise, low-entropy policies. Through experiments on a variety of control problems, we concretely demonstrate this misleading effect. Our analysis leads to better understanding of how to balance reward design and entropy maximization in challenging control problems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Maximum Entropy Misleads Policy Optimization
Zhang, Ruipeng
Chang, Ya-Chien
Gao, Sicun
Machine Learning
Artificial Intelligence
The Maximum Entropy Reinforcement Learning (MaxEnt RL) framework is a leading approach for achieving efficient learning and robust performance across many RL tasks. However, MaxEnt methods have also been shown to struggle with performance-critical control problems in practice, where non-MaxEnt algorithms can successfully learn. In this work, we analyze how the trade-off between robustness and optimality affects the performance of MaxEnt algorithms in complex control tasks: while entropy maximization enhances exploration and robustness, it can also mislead policy optimization, leading to failure in tasks that require precise, low-entropy policies. Through experiments on a variety of control problems, we concretely demonstrate this misleading effect. Our analysis leads to better understanding of how to balance reward design and entropy maximization in challenging control problems.
title When Maximum Entropy Misleads Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.05615