Average-Reward Soft Actor-Critic

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Adamczyk, Jacob, Makarenko, Volodymyr, Tiomkin, Stas, Kulkarni, Rahul V.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915427618127872
author Adamczyk, Jacob
Makarenko, Volodymyr
Tiomkin, Stas
Kulkarni, Rahul V.
author_facet Adamczyk, Jacob
Makarenko, Volodymyr
Tiomkin, Stas
Kulkarni, Rahul V.
contents The average-reward formulation of reinforcement learning (RL) has drawn increased interest in recent years for its ability to solve temporally-extended problems without relying on discounting. Meanwhile, in the discounted setting, algorithms with entropy regularization have been developed, leading to improvements over deterministic methods. Despite the distinct benefits of these approaches, deep RL algorithms for the entropy-regularized average-reward objective have not been developed. While policy-gradient based approaches have recently been presented for the average-reward literature, the corresponding actor-critic framework remains less explored. In this paper, we introduce an average-reward soft actor-critic algorithm to address these gaps in the field. We validate our method by comparing with existing average-reward algorithms on standard RL benchmarks, achieving superior performance for the average-reward criterion.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09080
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Average-Reward Soft Actor-Critic
Adamczyk, Jacob
Makarenko, Volodymyr
Tiomkin, Stas
Kulkarni, Rahul V.
Machine Learning
Artificial Intelligence
The average-reward formulation of reinforcement learning (RL) has drawn increased interest in recent years for its ability to solve temporally-extended problems without relying on discounting. Meanwhile, in the discounted setting, algorithms with entropy regularization have been developed, leading to improvements over deterministic methods. Despite the distinct benefits of these approaches, deep RL algorithms for the entropy-regularized average-reward objective have not been developed. While policy-gradient based approaches have recently been presented for the average-reward literature, the corresponding actor-critic framework remains less explored. In this paper, we introduce an average-reward soft actor-critic algorithm to address these gaps in the field. We validate our method by comparing with existing average-reward algorithms on standard RL benchmarks, achieving superior performance for the average-reward criterion.
title Average-Reward Soft Actor-Critic
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.09080