ADDQ: Adaptive Distributional Double Q-Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913911319560192 |
|---|---|
| author | Döring, Leif Wille, Benedikt Birr, Maximilian Bîrsan, Mihail Slowik, Martin |
| author_facet | Döring, Leif Wille, Benedikt Birr, Maximilian Bîrsan, Mihail Slowik, Martin |
| contents | Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direct or indirect overestimation reduction mechanism. We propose an easy to implement method built on top of distributional reinforcement learning (DRL) algorithms to deal with the overestimation in a locally adaptive way. Our framework is simple to implement, existing distributional algorithms can be improved with a few lines of code. We provide theoretical evidence and use double $Q$-learning to show how to include locally adaptive overestimation control in existing algorithms. Experiments are provided for tabular, Atari, and MuJoCo environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_19478 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ADDQ: Adaptive Distributional Double Q-Learning Döring, Leif Wille, Benedikt Birr, Maximilian Bîrsan, Mihail Slowik, Martin Machine Learning Optimization and Control Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direct or indirect overestimation reduction mechanism. We propose an easy to implement method built on top of distributional reinforcement learning (DRL) algorithms to deal with the overestimation in a locally adaptive way. Our framework is simple to implement, existing distributional algorithms can be improved with a few lines of code. We provide theoretical evidence and use double $Q$-learning to show how to include locally adaptive overestimation control in existing algorithms. Experiments are provided for tabular, Atari, and MuJoCo environments. |
| title | ADDQ: Adaptive Distributional Double Q-Learning |
| topic | Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2506.19478 |