DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Xiaoteng, Chen, Junyao, Xia, Li, Yang, Jun, Zhao, Qianchuan, Zhou, Zhengyuan
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911026646089728
author Ma, Xiaoteng
Chen, Junyao
Xia, Li
Yang, Jun
Zhao, Qianchuan
Zhou, Zhengyuan
author_facet Ma, Xiaoteng
Chen, Junyao
Xia, Li
Yang, Jun
Zhao, Qianchuan
Zhou, Zhengyuan
contents We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2004_14547
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
Ma, Xiaoteng
Chen, Junyao
Xia, Li
Yang, Jun
Zhao, Qianchuan
Zhou, Zhengyuan
Machine Learning
Artificial Intelligence
We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.
title DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2004.14547