Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2310.17173 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918066910134272 |
|---|---|
| author | Neo, Dexter Chen, Tsuhan |
| author_facet | Neo, Dexter Chen, Tsuhan |
| contents | We present a novel extension to the family of Soft Actor-Critic (SAC) algorithms. We argue that based on the Maximum Entropy Principle, discrete SAC can be further improved via additional statistical constraints derived from a surrogate critic policy. Furthermore, our findings suggests that these constraints provide an added robustness against potential domain shifts, which are essential for safe deployment of reinforcement learning agents in the real-world. We provide theoretical analysis and show empirical results on low data regimes for both in-distribution and out-of-distribution variants of Atari 2600 games. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_17173 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic Neo, Dexter Chen, Tsuhan Machine Learning We present a novel extension to the family of Soft Actor-Critic (SAC) algorithms. We argue that based on the Maximum Entropy Principle, discrete SAC can be further improved via additional statistical constraints derived from a surrogate critic policy. Furthermore, our findings suggests that these constraints provide an added robustness against potential domain shifts, which are essential for safe deployment of reinforcement learning agents in the real-world. We provide theoretical analysis and show empirical results on low data regimes for both in-distribution and out-of-distribution variants of Atari 2600 games. |
| title | DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2310.17173 |