Is Q-learning an Ill-posed Problem?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wissmann, Philipp, Hein, Daniel, Udluft, Steffen, Runkler, Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915165366124544
author Wissmann, Philipp
Hein, Daniel
Udluft, Steffen
Runkler, Thomas
author_facet Wissmann, Philipp
Hein, Daniel
Udluft, Steffen
Runkler, Thomas
contents This paper investigates the instability of Q-learning in continuous environments, a challenge frequently encountered by practitioners. Traditionally, this instability is attributed to bootstrapping and regression model errors. Using a representative reinforcement learning benchmark, we systematically examine the effects of bootstrapping and model inaccuracies by incrementally eliminating these potential error sources. Our findings reveal that even in relatively simple benchmarks, the fundamental task of Q-learning - iteratively learning a Q-function from policy-specific target values - can be inherently ill-posed and prone to failure. These insights cast doubt on the reliability of Q-learning as a universal solution for reinforcement learning problems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is Q-learning an Ill-posed Problem?
Wissmann, Philipp
Hein, Daniel
Udluft, Steffen
Runkler, Thomas
Machine Learning
Artificial Intelligence
This paper investigates the instability of Q-learning in continuous environments, a challenge frequently encountered by practitioners. Traditionally, this instability is attributed to bootstrapping and regression model errors. Using a representative reinforcement learning benchmark, we systematically examine the effects of bootstrapping and model inaccuracies by incrementally eliminating these potential error sources. Our findings reveal that even in relatively simple benchmarks, the fundamental task of Q-learning - iteratively learning a Q-function from policy-specific target values - can be inherently ill-posed and prone to failure. These insights cast doubt on the reliability of Q-learning as a universal solution for reinforcement learning problems.
title Is Q-learning an Ill-posed Problem?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.14365