PAC-Bayesian Soft Actor-Critic Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tasdighi, Bahareh, Akgül, Abdullah, Haussmann, Manuel, Brink, Kenny Kazimirzak, Kandemir, Melih
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917688950915072
author Tasdighi, Bahareh
Akgül, Abdullah
Haussmann, Manuel
Brink, Kenny Kazimirzak
Kandemir, Melih
author_facet Tasdighi, Bahareh
Akgül, Abdullah
Haussmann, Manuel
Brink, Kenny Kazimirzak
Kandemir, Melih
contents Actor-critic algorithms address the dual goals of reinforcement learning (RL), policy evaluation and improvement via two separate function approximators. The practicality of this approach comes at the expense of training instability, caused mainly by the destructive effect of the approximation errors of the critic on the actor. We tackle this bottleneck by employing an existing Probably Approximately Correct (PAC) Bayesian bound for the first time as the critic training objective of the Soft Actor-Critic (SAC) algorithm. We further demonstrate that online learning performance improves significantly when a stochastic actor explores multiple futures by critic-guided random search. We observe our resulting algorithm to compare favorably against the state-of-the-art SAC implementation on multiple classical control and locomotion tasks in terms of both sample efficiency and regret.
format Preprint
id arxiv_https___arxiv_org_abs_2301_12776
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle PAC-Bayesian Soft Actor-Critic Learning
Tasdighi, Bahareh
Akgül, Abdullah
Haussmann, Manuel
Brink, Kenny Kazimirzak
Kandemir, Melih
Machine Learning
Actor-critic algorithms address the dual goals of reinforcement learning (RL), policy evaluation and improvement via two separate function approximators. The practicality of this approach comes at the expense of training instability, caused mainly by the destructive effect of the approximation errors of the critic on the actor. We tackle this bottleneck by employing an existing Probably Approximately Correct (PAC) Bayesian bound for the first time as the critic training objective of the Soft Actor-Critic (SAC) algorithm. We further demonstrate that online learning performance improves significantly when a stochastic actor explores multiple futures by critic-guided random search. We observe our resulting algorithm to compare favorably against the state-of-the-art SAC implementation on multiple classical control and locomotion tasks in terms of both sample efficiency and regret.
title PAC-Bayesian Soft Actor-Critic Learning
topic Machine Learning
url https://arxiv.org/abs/2301.12776