ADDQ: Adaptive Distributional Double Q-Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Döring, Leif, Wille, Benedikt, Birr, Maximilian, Bîrsan, Mihail, Slowik, Martin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913911319560192
author Döring, Leif
Wille, Benedikt
Birr, Maximilian
Bîrsan, Mihail
Slowik, Martin
author_facet Döring, Leif
Wille, Benedikt
Birr, Maximilian
Bîrsan, Mihail
Slowik, Martin
contents Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direct or indirect overestimation reduction mechanism. We propose an easy to implement method built on top of distributional reinforcement learning (DRL) algorithms to deal with the overestimation in a locally adaptive way. Our framework is simple to implement, existing distributional algorithms can be improved with a few lines of code. We provide theoretical evidence and use double $Q$-learning to show how to include locally adaptive overestimation control in existing algorithms. Experiments are provided for tabular, Atari, and MuJoCo environments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ADDQ: Adaptive Distributional Double Q-Learning
Döring, Leif
Wille, Benedikt
Birr, Maximilian
Bîrsan, Mihail
Slowik, Martin
Machine Learning
Optimization and Control
Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direct or indirect overestimation reduction mechanism. We propose an easy to implement method built on top of distributional reinforcement learning (DRL) algorithms to deal with the overestimation in a locally adaptive way. Our framework is simple to implement, existing distributional algorithms can be improved with a few lines of code. We provide theoretical evidence and use double $Q$-learning to show how to include locally adaptive overestimation control in existing algorithms. Experiments are provided for tabular, Atari, and MuJoCo environments.
title ADDQ: Adaptive Distributional Double Q-Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2506.19478