DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914139253768192 |
|---|---|
| author | Perkins, Daniel Escobar, Oscar J. Green, Luke |
| author_facet | Perkins, Daniel Escobar, Oscar J. Green, Luke |
| contents | We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward optimization. We investigate how prioritized experience replay leads to faster convergence and higher returns and show empirical results comparing uniform, no replay, and prioritized strategies across multiple simulations. Our findings illuminate the trade-offs and interactions between exploration strategies and memory management in DQN training, offering practical recommendations for robust reinforcement learning in resource-constrained settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_03670 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay Perkins, Daniel Escobar, Oscar J. Green, Luke Machine Learning Artificial Intelligence 68T05 We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward optimization. We investigate how prioritized experience replay leads to faster convergence and higher returns and show empirical results comparing uniform, no replay, and prioritized strategies across multiple simulations. Our findings illuminate the trade-offs and interactions between exploration strategies and memory management in DQN training, offering practical recommendations for robust reinforcement learning in resource-constrained settings. |
| title | DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay |
| topic | Machine Learning Artificial Intelligence 68T05 |
| url | https://arxiv.org/abs/2511.03670 |