A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901404408348672 |
|---|---|
| author | MD Israfeel |
| author_facet | MD Israfeel |
| contents | <p>This paper presents a reinforcement learning calibration benchmark comparing tabular Q-Learning and one-hot Deep Q-Networks (DQN) on enumerable maze navigation tasks. Across deterministic grid-world environments with matched interaction budgets and multi-seed evaluation, tabular Q-Learning achieves equivalent final policy quality with superior sample efficiency and lower wall-clock cost. Additional ablations isolate the contribution of depth, replay buffers, target networks, and state representation. Results suggest that, in small enumerable MDPs, optimization overhead from deep RL infrastructure can dominate any benefit from function approximation. The repository includes a fully reproducible single-file NumPy implementation, statistical evaluation utilities, and benchmark scripts.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20075861 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation MD Israfeel deep reinforcement learning reinforcement learning DQN Q-learning gridworld maze navigation <p>This paper presents a reinforcement learning calibration benchmark comparing tabular Q-Learning and one-hot Deep Q-Networks (DQN) on enumerable maze navigation tasks. Across deterministic grid-world environments with matched interaction budgets and multi-seed evaluation, tabular Q-Learning achieves equivalent final policy quality with superior sample efficiency and lower wall-clock cost. Additional ablations isolate the contribution of depth, replay buffers, target networks, and state representation. Results suggest that, in small enumerable MDPs, optimization overhead from deep RL infrastructure can dominate any benefit from function approximation. The repository includes a fully reproducible single-file NumPy implementation, statistical evaluation utilities, and benchmark scripts.</p> |
| title | A Reinforcement Learning Calibration Benchmark of Tabular Q-Learning and One-Hot DQN on Enumerable Maze Navigation |
| topic | deep reinforcement learning reinforcement learning DQN Q-learning gridworld maze navigation |
| url | https://doi.org/10.5281/zenodo.20075861 |