Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
Fuente:
arXiv
Guardado en:
| Autores principales: | Lim, Han-Dong, Lee, Donghwan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A finite time analysis of distributed Q-learning
por: Lim, Han-Dong, et al.
Publicado: (2024)
por: Lim, Han-Dong, et al.
Publicado: (2024)
Periodic Regularized Q-Learning
por: Yang, Hyukjun, et al.
Publicado: (2026)
por: Yang, Hyukjun, et al.
Publicado: (2026)
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
por: Lim, Han-Dong, et al.
Publicado: (2024)
por: Lim, Han-Dong, et al.
Publicado: (2024)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
por: Lim, Han-Dong, et al.
Publicado: (2025)
por: Lim, Han-Dong, et al.
Publicado: (2025)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
por: Lim, Han-Dong, et al.
Publicado: (2023)
por: Lim, Han-Dong, et al.
Publicado: (2023)
Backstepping Temporal Difference Learning
por: Lim, Han-Dong, et al.
Publicado: (2023)
por: Lim, Han-Dong, et al.
Publicado: (2023)
Safe-Support Q-Learning: Learning without Unsafe Exploration
por: Lim, Yeeun, et al.
Publicado: (2026)
por: Lim, Yeeun, et al.
Publicado: (2026)
Regularized Q-learning
por: Lim, Han-Dong, et al.
Publicado: (2022)
por: Lim, Han-Dong, et al.
Publicado: (2022)
Lyapunov-Certified Direct Switching Theory for Q-Learning
por: Lee, Donghwan
Publicado: (2026)
por: Lee, Donghwan
Publicado: (2026)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
por: Lee, HyeAnn, et al.
Publicado: (2023)
por: Lee, HyeAnn, et al.
Publicado: (2023)
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
por: Lee, Donghwan, et al.
Publicado: (2024)
por: Lee, Donghwan, et al.
Publicado: (2024)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
por: Park, Jongchan, et al.
Publicado: (2025)
por: Park, Jongchan, et al.
Publicado: (2025)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2024)
por: Omura, Motoki, et al.
Publicado: (2024)
Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem
por: Yang, Hyukjun, et al.
Publicado: (2026)
por: Yang, Hyukjun, et al.
Publicado: (2026)
A Switching System Theory of Q-Learning with Linear Function Approximation
por: Lee, Donghwan, et al.
Publicado: (2026)
por: Lee, Donghwan, et al.
Publicado: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
por: Lee, Taeho, et al.
Publicado: (2026)
por: Lee, Taeho, et al.
Publicado: (2026)
Adaptive Policy Backbone via Shared Network
por: Park, Bumgeun, et al.
Publicado: (2025)
por: Park, Bumgeun, et al.
Publicado: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
por: Na, Hyunjun, et al.
Publicado: (2026)
por: Na, Hyunjun, et al.
Publicado: (2026)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
por: Na, Hyunjun, et al.
Publicado: (2026)
por: Na, Hyunjun, et al.
Publicado: (2026)
Bellman operator convergence enhancements in reinforcement learning algorithms
por: Kadurha, David Krame, et al.
Publicado: (2025)
por: Kadurha, David Krame, et al.
Publicado: (2025)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2024)
por: Vincent, Théo, et al.
Publicado: (2024)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
por: Cho, Taehyun, et al.
Publicado: (2024)
por: Cho, Taehyun, et al.
Publicado: (2024)
A primal-dual perspective for distributed TD-learning
por: Lim, Han-Dong, et al.
Publicado: (2023)
por: Lim, Han-Dong, et al.
Publicado: (2023)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
por: Yeom, Junghyuk, et al.
Publicado: (2024)
por: Yeom, Junghyuk, et al.
Publicado: (2024)
Bellman Error Centering
por: Chen, Xingguo, et al.
Publicado: (2025)
por: Chen, Xingguo, et al.
Publicado: (2025)
Parameterized Projected Bellman Operator
por: Vincent, Théo, et al.
Publicado: (2023)
por: Vincent, Théo, et al.
Publicado: (2023)
An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
por: Xu, Wei, et al.
Publicado: (2025)
por: Xu, Wei, et al.
Publicado: (2025)
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
por: Lee, Donghwan
Publicado: (2026)
por: Lee, Donghwan
Publicado: (2026)
MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse
por: Kim, Donghwan, et al.
Publicado: (2026)
por: Kim, Donghwan, et al.
Publicado: (2026)
Mitigating the Likelihood Paradox in Flow-based OOD Detection via Entropy Manipulation
por: Kim, Donghwan, et al.
Publicado: (2026)
por: Kim, Donghwan, et al.
Publicado: (2026)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2025)
por: Omura, Motoki, et al.
Publicado: (2025)
Theoretical Barriers in Bellman-Based Reinforcement Learning
por: Pinon, Brieuc, et al.
Publicado: (2025)
por: Pinon, Brieuc, et al.
Publicado: (2025)
Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence
por: Lee, Donghwan, et al.
Publicado: (2026)
por: Lee, Donghwan, et al.
Publicado: (2026)
Adversarial bandit optimization for approximately linear functions
por: Cheng, Zhuoyu, et al.
Publicado: (2025)
por: Cheng, Zhuoyu, et al.
Publicado: (2025)
Deep Double Q-learning
por: Nagarajan, Prabhat, et al.
Publicado: (2025)
por: Nagarajan, Prabhat, et al.
Publicado: (2025)
Why the Counterintuitive Phenomenon of Likelihood Rarely Appears in Tabular Anomaly Detection with Deep Generative Models?
por: Kim, Donghwan, et al.
Publicado: (2026)
por: Kim, Donghwan, et al.
Publicado: (2026)
Adaptive Regularization of Representation Rank as an Implicit Constraint of Bellman Equation
por: He, Qiang, et al.
Publicado: (2024)
por: He, Qiang, et al.
Publicado: (2024)
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
por: Lee, Dohyeok, et al.
Publicado: (2024)
por: Lee, Dohyeok, et al.
Publicado: (2024)
Is Q-learning an Ill-posed Problem?
por: Wissmann, Philipp, et al.
Publicado: (2025)
por: Wissmann, Philipp, et al.
Publicado: (2025)
MinMaxMin $Q$-learning
por: Soffair, Nitsan, et al.
Publicado: (2024)
por: Soffair, Nitsan, et al.
Publicado: (2024)
Ejemplares similares
-
A finite time analysis of distributed Q-learning
por: Lim, Han-Dong, et al.
Publicado: (2024) -
Periodic Regularized Q-Learning
por: Yang, Hyukjun, et al.
Publicado: (2026) -
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
por: Lim, Han-Dong, et al.
Publicado: (2024) -
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
por: Lim, Han-Dong, et al.
Publicado: (2025) -
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
por: Lim, Han-Dong, et al.
Publicado: (2023)