Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913309607854080 |
|---|---|
| author | Rahn, Nate D'Oro, Pierluca Wiltzer, Harley Bacon, Pierre-Luc Bellemare, Marc G. |
| author_facet | Rahn, Nate D'Oro, Pierluca Wiltzer, Harley Bacon, Pierre-Luc Bellemare, Marc G. |
| contents | Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective on these behaviors by studying the return landscape: the mapping between a policy and a return. We find that popular algorithms traverse noisy neighborhoods of this landscape, in which a single update to the policy parameters leads to a wide range of returns. By taking a distributional view of these returns, we map the landscape, characterizing failure-prone regions of policy space and revealing a hidden dimension of policy quality. We show that the landscape exhibits surprising structure by finding simple paths in parameter space which improve the stability of a policy. To conclude, we develop a distribution-aware procedure which finds such paths, navigating away from noisy neighborhoods in order to improve the robustness of a policy. Taken together, our results provide new insight into the optimization, evaluation, and design of agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_14597 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control Rahn, Nate D'Oro, Pierluca Wiltzer, Harley Bacon, Pierre-Luc Bellemare, Marc G. Machine Learning Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective on these behaviors by studying the return landscape: the mapping between a policy and a return. We find that popular algorithms traverse noisy neighborhoods of this landscape, in which a single update to the policy parameters leads to a wide range of returns. By taking a distributional view of these returns, we map the landscape, characterizing failure-prone regions of policy space and revealing a hidden dimension of policy quality. We show that the landscape exhibits surprising structure by finding simple paths in parameter space which improve the stability of a policy. To conclude, we develop a distribution-aware procedure which finds such paths, navigating away from noisy neighborhoods in order to improve the robustness of a policy. Taken together, our results provide new insight into the optimization, evaluation, and design of agents. |
| title | Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2309.14597 |