Is Q-learning an Ill-posed Problem?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915165366124544 |
|---|---|
| author | Wissmann, Philipp Hein, Daniel Udluft, Steffen Runkler, Thomas |
| author_facet | Wissmann, Philipp Hein, Daniel Udluft, Steffen Runkler, Thomas |
| contents | This paper investigates the instability of Q-learning in continuous environments, a challenge frequently encountered by practitioners. Traditionally, this instability is attributed to bootstrapping and regression model errors. Using a representative reinforcement learning benchmark, we systematically examine the effects of bootstrapping and model inaccuracies by incrementally eliminating these potential error sources. Our findings reveal that even in relatively simple benchmarks, the fundamental task of Q-learning - iteratively learning a Q-function from policy-specific target values - can be inherently ill-posed and prone to failure. These insights cast doubt on the reliability of Q-learning as a universal solution for reinforcement learning problems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_14365 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Is Q-learning an Ill-posed Problem? Wissmann, Philipp Hein, Daniel Udluft, Steffen Runkler, Thomas Machine Learning Artificial Intelligence This paper investigates the instability of Q-learning in continuous environments, a challenge frequently encountered by practitioners. Traditionally, this instability is attributed to bootstrapping and regression model errors. Using a representative reinforcement learning benchmark, we systematically examine the effects of bootstrapping and model inaccuracies by incrementally eliminating these potential error sources. Our findings reveal that even in relatively simple benchmarks, the fundamental task of Q-learning - iteratively learning a Q-function from policy-specific target values - can be inherently ill-posed and prone to failure. These insights cast doubt on the reliability of Q-learning as a universal solution for reinforcement learning problems. |
| title | Is Q-learning an Ill-posed Problem? |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2502.14365 |