Evolutionary Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mustafaoglu, Zelal Su "Lain", Pingali, Keshav, Miikkulainen, Risto
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909582236844032
author Mustafaoglu, Zelal Su "Lain"
Pingali, Keshav
Miikkulainen, Risto
author_facet Mustafaoglu, Zelal Su "Lain"
Pingali, Keshav
Miikkulainen, Risto
contents A key challenge in reinforcement learning (RL) is managing the exploration-exploitation trade-off without sacrificing sample efficiency. Policy gradient (PG) methods excel in exploitation through fine-grained, gradient-based optimization but often struggle with exploration due to their focus on local search. In contrast, evolutionary computation (EC) methods excel in global exploration, but lack mechanisms for exploitation. To address these limitations, this paper proposes Evolutionary Policy Optimization (EPO), a hybrid algorithm that integrates neuroevolution with policy gradient methods for policy optimization. EPO leverages the exploration capabilities of EC and the exploitation strengths of PG, offering an efficient solution to the exploration-exploitation dilemma in RL. EPO is evaluated on the Atari Pong and Breakout benchmarks. Experimental results show that EPO improves both policy quality and sample efficiency compared to standard PG and EC methods, making it effective for tasks that require both exploration and local optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12568
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolutionary Policy Optimization
Mustafaoglu, Zelal Su "Lain"
Pingali, Keshav
Miikkulainen, Risto
Machine Learning
Neural and Evolutionary Computing
A key challenge in reinforcement learning (RL) is managing the exploration-exploitation trade-off without sacrificing sample efficiency. Policy gradient (PG) methods excel in exploitation through fine-grained, gradient-based optimization but often struggle with exploration due to their focus on local search. In contrast, evolutionary computation (EC) methods excel in global exploration, but lack mechanisms for exploitation. To address these limitations, this paper proposes Evolutionary Policy Optimization (EPO), a hybrid algorithm that integrates neuroevolution with policy gradient methods for policy optimization. EPO leverages the exploration capabilities of EC and the exploitation strengths of PG, offering an efficient solution to the exploration-exploitation dilemma in RL. EPO is evaluated on the Atari Pong and Breakout benchmarks. Experimental results show that EPO improves both policy quality and sample efficiency compared to standard PG and EC methods, making it effective for tasks that require both exploration and local optimization.
title Evolutionary Policy Optimization
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2504.12568