DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2020
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911026646089728 |
|---|---|
| author | Ma, Xiaoteng Chen, Junyao Xia, Li Yang, Jun Zhao, Qianchuan Zhou, Zhengyuan |
| author_facet | Ma, Xiaoteng Chen, Junyao Xia, Li Yang, Jun Zhao, Qianchuan Zhou, Zhengyuan |
| contents | We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2004_14547 |
| institution | arXiv |
| publishDate | 2020 |
| record_format | arxiv |
| spellingShingle | DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning Ma, Xiaoteng Chen, Junyao Xia, Li Yang, Jun Zhao, Qianchuan Zhou, Zhengyuan Machine Learning Artificial Intelligence We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks. |
| title | DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2004.14547 |