A Switching System Theory of Q-Learning with Linear Function Approximation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Donghwan, Lim, Han-Dong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918509996408832
author Lee, Donghwan
Lim, Han-Dong
author_facet Lee, Donghwan
Lim, Han-Dong
contents This paper develops a switching-system interpretation of Q-learning with linear function approximation (LFA) based on the joint spectral radius (JSR). We derive an exact linear switched model for the mean dynamics and relate convergence to stability of the corresponding switched system. The same construction is then used for stochastic linear Q-learning with independent and identically distributed (i.i.d.) observations and with Markovian observations. Although exact JSR computation is difficult in general, the certificate captures products of switching modes and can be less conservative than one-step norm bounds. The framework also yields a JSR-based view of regularized Q-learning with LFA. The resulting analysis connects projected Bellman equations, finite-difference stochastic-policy switching, and switched-system stability in a single parameter-space formulation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Switching System Theory of Q-Learning with Linear Function Approximation
Lee, Donghwan
Lim, Han-Dong
Machine Learning
This paper develops a switching-system interpretation of Q-learning with linear function approximation (LFA) based on the joint spectral radius (JSR). We derive an exact linear switched model for the mean dynamics and relate convergence to stability of the corresponding switched system. The same construction is then used for stochastic linear Q-learning with independent and identically distributed (i.i.d.) observations and with Markovian observations. Although exact JSR computation is difficult in general, the certificate captures products of switching modes and can be less conservative than one-step norm bounds. The framework also yields a JSR-based view of regularized Q-learning with LFA. The resulting analysis connects projected Bellman equations, finite-difference stochastic-policy switching, and switched-system stability in a single parameter-space formulation.
title A Switching System Theory of Q-Learning with Linear Function Approximation
topic Machine Learning
url https://arxiv.org/abs/2605.11021