Is Q-learning an Ill-posed Problem?
Fuente:
arXiv
Saved in:
| Main Authors: | Wissmann, Philipp, Hein, Daniel, Udluft, Steffen, Runkler, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model-based Offline Quantum Reinforcement Learning
by: Eisenmann, Simon, et al.
Published: (2024)
by: Eisenmann, Simon, et al.
Published: (2024)
Why long model-based rollouts are no reason for bad Q-value estimates
by: Wissmann, Philipp, et al.
Published: (2024)
by: Wissmann, Philipp, et al.
Published: (2024)
Variational Quantum Circuits in Offline Contextual Bandit Problems
by: Schulte, Lukas, et al.
Published: (2025)
by: Schulte, Lukas, et al.
Published: (2025)
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning
by: Ormancı, Batıkan Bora, et al.
Published: (2024)
by: Ormancı, Batıkan Bora, et al.
Published: (2024)
Learning Control Policies for Variable Objectives from Offline Data
by: Weber, Marc, et al.
Published: (2023)
by: Weber, Marc, et al.
Published: (2023)
Deep Double Q-learning
by: Nagarajan, Prabhat, et al.
Published: (2025)
by: Nagarajan, Prabhat, et al.
Published: (2025)
On-device Online Learning and Semantic Management of TinyML Systems
by: Ren, Haoyu, et al.
Published: (2024)
by: Ren, Haoyu, et al.
Published: (2024)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
by: Röpke, Willem, et al.
Published: (2026)
by: Röpke, Willem, et al.
Published: (2026)
MinMaxMin $Q$-learning
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Stabilizing Extreme Q-learning by Maclaurin Expansion
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Q-learning with Adjoint Matching
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
Neural-ANOVA: Analytical Model Decomposition using Automatic Integration
by: Limmer, Steffen, et al.
Published: (2024)
by: Limmer, Steffen, et al.
Published: (2024)
Comprehensive Evaluation of Prototype Neural Networks
by: Schlinge, Philipp, et al.
Published: (2025)
by: Schlinge, Philipp, et al.
Published: (2025)
Yes, Q-learning Helps Offline In-Context RL
by: Tarasov, Denis, et al.
Published: (2025)
by: Tarasov, Denis, et al.
Published: (2025)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
by: Yeom, Junghyuk, et al.
Published: (2024)
by: Yeom, Junghyuk, et al.
Published: (2024)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Compositional Q-learning for electrolyte repletion with imbalanced patient sub-populations
by: Mandyam, Aishwarya, et al.
Published: (2021)
by: Mandyam, Aishwarya, et al.
Published: (2021)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Scaling CrossQ with Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously Thought
by: Yan, Cheng, et al.
Published: (2025)
by: Yan, Cheng, et al.
Published: (2025)
SortBench: Benchmarking LLMs based on their ability to sort lists
by: Herbold, Steffen
Published: (2025)
by: Herbold, Steffen
Published: (2025)
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
by: Meng, Li, et al.
Published: (2021)
by: Meng, Li, et al.
Published: (2021)
Stochastic Q-learning for Large Discrete Action Spaces
by: Fourati, Fares, et al.
Published: (2024)
by: Fourati, Fares, et al.
Published: (2024)
A finite time analysis of distributed Q-learning
by: Lim, Han-Dong, et al.
Published: (2024)
by: Lim, Han-Dong, et al.
Published: (2024)
Towards a Problem-Oriented Domain Adaptation Framework for Machine Learning
by: Spitzer, Philipp, et al.
Published: (2025)
by: Spitzer, Philipp, et al.
Published: (2025)
LinearizeLLM: An Agent-Based Framework for LLM-Driven Exact Linear Reformulation of Nonlinear Optimization Problems
by: Kandora, Paul-Niklas Ken, et al.
Published: (2025)
by: Kandora, Paul-Niklas Ken, et al.
Published: (2025)
Probing Implicit Bias in Semi-gradient Q-learning: Visualizing the Effective Loss Landscapes via the Fokker--Planck Equation
by: Yin, Shuyu, et al.
Published: (2024)
by: Yin, Shuyu, et al.
Published: (2024)
Deep Reinforcement Learning with Spiking Q-learning
by: Chen, Ding, et al.
Published: (2022)
by: Chen, Ding, et al.
Published: (2022)
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
Frictional Q-Learning
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Flow Q-Learning
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Drift Q-Learning
by: Houssaini, Anas, et al.
Published: (2026)
by: Houssaini, Anas, et al.
Published: (2026)
Cost-optimal Sequential Testing via Doubly Robust Q-learning
by: Zhou, Doudou, et al.
Published: (2026)
by: Zhou, Doudou, et al.
Published: (2026)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
by: Tang, Yongjian, et al.
Published: (2026)
by: Tang, Yongjian, et al.
Published: (2026)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Scalable In-Context Q-Learning
by: Liu, Jinmei, et al.
Published: (2025)
by: Liu, Jinmei, et al.
Published: (2025)
Chunk-Guided Q-Learning
by: Song, Gwanwoo, et al.
Published: (2026)
by: Song, Gwanwoo, et al.
Published: (2026)
Similar Items
-
Model-based Offline Quantum Reinforcement Learning
by: Eisenmann, Simon, et al.
Published: (2024) -
Why long model-based rollouts are no reason for bad Q-value estimates
by: Wissmann, Philipp, et al.
Published: (2024) -
Variational Quantum Circuits in Offline Contextual Bandit Problems
by: Schulte, Lukas, et al.
Published: (2025) -
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning
by: Ormancı, Batıkan Bora, et al.
Published: (2024) -
Learning Control Policies for Variable Objectives from Offline Data
by: Weber, Marc, et al.
Published: (2023)