Neighboring State-based Exploration for Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917053790683136 |
|---|---|
| author | Li, Yu-Teng Lin, Justin Cheng, Jeffery Pachuca, Pedro |
| author_facet | Li, Yu-Teng Lin, Justin Cheng, Jeffery Pachuca, Pedro |
| contents | Reinforcement Learning is a powerful tool to model decision-making processes. However, it relies on an exploration-exploitation trade-off that remains an open challenge for many tasks. In this work, we study neighboring state-based, model-free exploration led by the intuition that, for an early-stage agent, considering actions derived from a bounded region of nearby states may lead to better actions when exploring. We propose two algorithms that choose exploratory actions based on a survey of nearby states, and find that one of our methods, $ρ$-explore, consistently outperforms the Double DQN baseline in an discrete environment by 49% in terms of Eval Reward Return. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2212_10712 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Neighboring State-based Exploration for Reinforcement Learning Li, Yu-Teng Lin, Justin Cheng, Jeffery Pachuca, Pedro Machine Learning Artificial Intelligence Reinforcement Learning is a powerful tool to model decision-making processes. However, it relies on an exploration-exploitation trade-off that remains an open challenge for many tasks. In this work, we study neighboring state-based, model-free exploration led by the intuition that, for an early-stage agent, considering actions derived from a bounded region of nearby states may lead to better actions when exploring. We propose two algorithms that choose exploratory actions based on a survey of nearby states, and find that one of our methods, $ρ$-explore, consistently outperforms the Double DQN baseline in an discrete environment by 49% in terms of Eval Reward Return. |
| title | Neighboring State-based Exploration for Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2212.10712 |