A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: MD Israfeel
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901404408348672
author MD Israfeel
author_facet MD Israfeel
contents <p>This paper presents a reinforcement learning calibration benchmark comparing tabular Q-Learning and one-hot Deep Q-Networks (DQN) on enumerable maze navigation tasks. Across deterministic grid-world environments with matched interaction budgets and multi-seed evaluation, tabular Q-Learning achieves equivalent final policy quality with superior sample efficiency and lower wall-clock cost. Additional ablations isolate the contribution of depth, replay buffers, target networks, and state representation. Results suggest that, in small enumerable MDPs, optimization overhead from deep RL infrastructure can dominate any benefit from function approximation. The repository includes a fully reproducible single-file NumPy implementation, statistical evaluation utilities, and benchmark scripts.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20075861
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation
MD Israfeel
deep reinforcement learning
reinforcement learning
DQN
Q-learning
gridworld
maze navigation
<p>This paper presents a reinforcement learning calibration benchmark comparing tabular Q-Learning and one-hot Deep Q-Networks (DQN) on enumerable maze navigation tasks. Across deterministic grid-world environments with matched interaction budgets and multi-seed evaluation, tabular Q-Learning achieves equivalent final policy quality with superior sample efficiency and lower wall-clock cost. Additional ablations isolate the contribution of depth, replay buffers, target networks, and state representation. Results suggest that, in small enumerable MDPs, optimization overhead from deep RL infrastructure can dominate any benefit from function approximation. The repository includes a fully reproducible single-file NumPy implementation, statistical evaluation utilities, and benchmark scripts.</p>
title A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation
topic deep reinforcement learning
reinforcement learning
DQN
Q-learning
gridworld
maze navigation
url https://doi.org/10.5281/zenodo.20075861