DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Perkins, Daniel, Escobar, Oscar J., Green, Luke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914139253768192
author Perkins, Daniel
Escobar, Oscar J.
Green, Luke
author_facet Perkins, Daniel
Escobar, Oscar J.
Green, Luke
contents We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward optimization. We investigate how prioritized experience replay leads to faster convergence and higher returns and show empirical results comparing uniform, no replay, and prioritized strategies across multiple simulations. Our findings illuminate the trade-offs and interactions between exploration strategies and memory management in DQN training, offering practical recommendations for robust reinforcement learning in resource-constrained settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
Perkins, Daniel
Escobar, Oscar J.
Green, Luke
Machine Learning
Artificial Intelligence
68T05
We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward optimization. We investigate how prioritized experience replay leads to faster convergence and higher returns and show empirical results comparing uniform, no replay, and prioritized strategies across multiple simulations. Our findings illuminate the trade-offs and interactions between exploration strategies and memory management in DQN training, offering practical recommendations for robust reinforcement learning in resource-constrained settings.
title DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
topic Machine Learning
Artificial Intelligence
68T05
url https://arxiv.org/abs/2511.03670