A Distributional Analogue to the Successor Representation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909209668354048 |
|---|---|
| author | Wiltzer, Harley Farebrother, Jesse Gretton, Arthur Tang, Yunhao Barreto, André Dabney, Will Bellemare, Marc G. Rowland, Mark |
| author_facet | Wiltzer, Harley Farebrother, Jesse Gretton, Arthur Tang, Yunhao Barreto, André Dabney, Will Bellemare, Marc G. Rowland, Mark |
| contents | This paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process. Analogous to how the successor representation (SR) describes the expected consequences of behaving according to a given policy, our distributional successor measure (SM) describes the distributional consequences of this behaviour. We formulate the distributional SM as a distribution over distributions and provide theory connecting it with distributional and model-based reinforcement learning. Moreover, we propose an algorithm that learns the distributional SM from data by minimizing a two-level maximum mean discrepancy. Key to our method are a number of algorithmic techniques that are independently valuable for learning generative models of state. As an illustration of the usefulness of the distributional SM, we show that it enables zero-shot risk-sensitive policy evaluation in a way that was not previously possible. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_08530 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Distributional Analogue to the Successor Representation Wiltzer, Harley Farebrother, Jesse Gretton, Arthur Tang, Yunhao Barreto, André Dabney, Will Bellemare, Marc G. Rowland, Mark Machine Learning Artificial Intelligence This paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process. Analogous to how the successor representation (SR) describes the expected consequences of behaving according to a given policy, our distributional successor measure (SM) describes the distributional consequences of this behaviour. We formulate the distributional SM as a distribution over distributions and provide theory connecting it with distributional and model-based reinforcement learning. Moreover, we propose an algorithm that learns the distributional SM from data by minimizing a two-level maximum mean discrepancy. Key to our method are a number of algorithmic techniques that are independently valuable for learning generative models of state. As an illustration of the usefulness of the distributional SM, we show that it enables zero-shot risk-sensitive policy evaluation in a way that was not previously possible. |
| title | A Distributional Analogue to the Successor Representation |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2402.08530 |