Regularized Q-learning through Robust Averaging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmitt-Förster, Peter, Sutter, Tobias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910461570580480
author Schmitt-Förster, Peter
Sutter, Tobias
author_facet Schmitt-Förster, Peter
Sutter, Tobias
contents We propose a new Q-learning variant, called 2RA Q-learning, that addresses some weaknesses of existing Q-learning methods in a principled manner. One such weakness is an underlying estimation bias which cannot be controlled and often results in poor performance. We propose a distributionally robust estimator for the maximum expected value term, which allows us to precisely control the level of estimation bias introduced. The distributionally robust estimator admits a closed-form solution such that the proposed algorithm has a computational cost per iteration comparable to Watkins' Q-learning. For the tabular case, we show that 2RA Q-learning converges to the optimal policy and analyze its asymptotic mean-squared error. Lastly, we conduct numerical experiments for various settings, which corroborate our theoretical findings and indicate that 2RA Q-learning often performs better than existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02201
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Regularized Q-learning through Robust Averaging
Schmitt-Förster, Peter
Sutter, Tobias
Optimization and Control
Machine Learning
We propose a new Q-learning variant, called 2RA Q-learning, that addresses some weaknesses of existing Q-learning methods in a principled manner. One such weakness is an underlying estimation bias which cannot be controlled and often results in poor performance. We propose a distributionally robust estimator for the maximum expected value term, which allows us to precisely control the level of estimation bias introduced. The distributionally robust estimator admits a closed-form solution such that the proposed algorithm has a computational cost per iteration comparable to Watkins' Q-learning. For the tabular case, we show that 2RA Q-learning converges to the optimal policy and analyze its asymptotic mean-squared error. Lastly, we conduct numerical experiments for various settings, which corroborate our theoretical findings and indicate that 2RA Q-learning often performs better than existing methods.
title Regularized Q-learning through Robust Averaging
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2405.02201