Periodic Regularized Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Hyukjun, Lim, Han-Dong, Lee, Donghwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914303995543552
author Yang, Hyukjun
Lim, Han-Dong
Lee, Donghwan
author_facet Yang, Hyukjun
Lim, Han-Dong
Lee, Donghwan
contents In reinforcement learning (RL), Q-learning is a fundamental algorithm whose convergence is guaranteed in the tabular setting. However, this convergence guarantee does not hold under linear function approximation. To overcome this limitation, a significant line of research has introduced regularization techniques to ensure stable convergence under function approximation. In this work, we propose a new algorithm, periodic regularized Q-learning (PRQ). We first introduce regularization at the level of the projection operator and explicitly construct a regularized projected value iteration (RP-VI), subsequently extending it to a sample-based RL algorithm. By appropriately regularizing the projection operator, the resulting projected value iteration becomes a contraction. By extending this regularized projection into the stochastic setting, we establish the PRQ algorithm and provide a rigorous theoretical analysis that proves finite-time convergence guarantees for PRQ under linear function approximation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03301
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Periodic Regularized Q-Learning
Yang, Hyukjun
Lim, Han-Dong
Lee, Donghwan
Machine Learning
Artificial Intelligence
In reinforcement learning (RL), Q-learning is a fundamental algorithm whose convergence is guaranteed in the tabular setting. However, this convergence guarantee does not hold under linear function approximation. To overcome this limitation, a significant line of research has introduced regularization techniques to ensure stable convergence under function approximation. In this work, we propose a new algorithm, periodic regularized Q-learning (PRQ). We first introduce regularization at the level of the projection operator and explicitly construct a regularized projected value iteration (RP-VI), subsequently extending it to a sample-based RL algorithm. By appropriately regularizing the projection operator, the resulting projected value iteration becomes a contraction. By extending this regularized projection into the stochastic setting, we establish the PRQ algorithm and provide a rigorous theoretical analysis that proves finite-time convergence guarantees for PRQ under linear function approximation.
title Periodic Regularized Q-Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03301