Extremum-Seeking Action Selection for Accelerating Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Ya-Chien, Gao, Sicun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916189034250240
author Chang, Ya-Chien
Gao, Sicun
author_facet Chang, Ya-Chien
Gao, Sicun
contents Reinforcement learning for control over continuous spaces typically uses high-entropy stochastic policies, such as Gaussian distributions, for local exploration and estimating policy gradient to optimize performance. Many robotic control problems deal with complex unstable dynamics, where applying actions that are off the feasible control manifolds can quickly lead to undesirable divergence. In such cases, most samples taken from the ambient action space generate low-value trajectories that hardly contribute to policy improvement, resulting in slow or failed learning. We propose to improve action selection in this model-free RL setting by introducing additional adaptive control steps based on Extremum-Seeking Control (ESC). On each action sampled from stochastic policies, we apply sinusoidal perturbations and query for estimated Q-values as the response signal. Based on ESC, we then dynamically improve the sampled actions to be closer to nearby optima before applying them to the environment. Our methods can be easily added in standard policy optimization to improve learning efficiency, which we demonstrate in various control learning environments.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01598
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extremum-Seeking Action Selection for Accelerating Policy Optimization
Chang, Ya-Chien
Gao, Sicun
Machine Learning
Artificial Intelligence
Robotics
Reinforcement learning for control over continuous spaces typically uses high-entropy stochastic policies, such as Gaussian distributions, for local exploration and estimating policy gradient to optimize performance. Many robotic control problems deal with complex unstable dynamics, where applying actions that are off the feasible control manifolds can quickly lead to undesirable divergence. In such cases, most samples taken from the ambient action space generate low-value trajectories that hardly contribute to policy improvement, resulting in slow or failed learning. We propose to improve action selection in this model-free RL setting by introducing additional adaptive control steps based on Extremum-Seeking Control (ESC). On each action sampled from stochastic policies, we apply sinusoidal perturbations and query for estimated Q-values as the response signal. Based on ESC, we then dynamically improve the sampled actions to be closer to nearby optima before applying them to the environment. Our methods can be easily added in standard policy optimization to improve learning efficiency, which we demonstrate in various control learning environments.
title Extremum-Seeking Action Selection for Accelerating Policy Optimization
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2404.01598