Sign-Separated Finite-Time Error Analysis of Q-Learning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Lee, Donghwan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
Switching-Geometry Analysis of Deflated Q-Value Iteration
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
Lyapunov-Certified Direct Switching Theory for Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
A Discrete-Time Switching System Analysis of Q-learning
von: Lee, Donghwan, et al.
Veröffentlicht: (2021)
von: Lee, Donghwan, et al.
Veröffentlicht: (2021)
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
von: Lee, Donghwan, et al.
Veröffentlicht: (2024)
von: Lee, Donghwan, et al.
Veröffentlicht: (2024)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
von: Lee, HyeAnn, et al.
Veröffentlicht: (2023)
von: Lee, HyeAnn, et al.
Veröffentlicht: (2023)
Finite-Time Error Analysis of Soft Q-Learning: Switching System Approach
von: Jeong, Narim, et al.
Veröffentlicht: (2024)
von: Jeong, Narim, et al.
Veröffentlicht: (2024)
Periodic Regularized Q-Learning
von: Yang, Hyukjun, et al.
Veröffentlicht: (2026)
von: Yang, Hyukjun, et al.
Veröffentlicht: (2026)
Safe-Support Q-Learning: Learning without Unsafe Exploration
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
A finite time analysis of distributed Q-learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Simultaneous Double Q-learning
von: Na, Hyunjun, et al.
Veröffentlicht: (2024)
von: Na, Hyunjun, et al.
Veröffentlicht: (2024)
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
Robust Deterministic Policy Gradient for Disturbance Attenuation and Its Application to Quadrotor Control
von: Lee, Taeho, et al.
Veröffentlicht: (2025)
von: Lee, Taeho, et al.
Veröffentlicht: (2025)
Backstepping Temporal Difference Learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
von: Lee, Taeho, et al.
Veröffentlicht: (2026)
von: Lee, Taeho, et al.
Veröffentlicht: (2026)
Analysis of approximate linear programming solution to Markov decision problem with log barrier function
von: Lee, Donghwan, et al.
Veröffentlicht: (2025)
von: Lee, Donghwan, et al.
Veröffentlicht: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Adaptive Policy Backbone via Shared Network
von: Park, Bumgeun, et al.
Veröffentlicht: (2025)
von: Park, Bumgeun, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
von: Jeong, Narim, et al.
Veröffentlicht: (2026)
von: Jeong, Narim, et al.
Veröffentlicht: (2026)
Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes
von: Lee, Donghwan, et al.
Veröffentlicht: (2023)
von: Lee, Donghwan, et al.
Veröffentlicht: (2023)
MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
Mitigating the Likelihood Paradox in Flow-based OOD Detection via Entropy Manipulation
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
Q-Learning under Finite Model Uncertainty
von: Sester, Julian, et al.
Veröffentlicht: (2024)
von: Sester, Julian, et al.
Veröffentlicht: (2024)
Inference of Deterministic Finite Automata via Q-Learning
von: Hosseinkhani, Elaheh, et al.
Veröffentlicht: (2025)
von: Hosseinkhani, Elaheh, et al.
Veröffentlicht: (2025)
Frictional Q-Learning
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Finite-Time Analysis of MCTS in Continuous POMDP Planning
von: Kong, Da, et al.
Veröffentlicht: (2026)
von: Kong, Da, et al.
Veröffentlicht: (2026)
Chunk-Guided Q-Learning
von: Song, Gwanwoo, et al.
Veröffentlicht: (2026)
von: Song, Gwanwoo, et al.
Veröffentlicht: (2026)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
Merge and Bound: Direct Manipulations on Weights for Class Incremental Learning
von: Kim, Taehoon, et al.
Veröffentlicht: (2025)
von: Kim, Taehoon, et al.
Veröffentlicht: (2025)
Why the Counterintuitive Phenomenon of Likelihood Rarely Appears in Tabular Anomaly Detection with Deep Generative Models?
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
von: Lee, Jihyun, et al.
Veröffentlicht: (2026)
von: Lee, Jihyun, et al.
Veröffentlicht: (2026)
Find A Winning Sign: Sign Is All We Need to Win the Lottery
von: Oh, Junghun, et al.
Veröffentlicht: (2025)
von: Oh, Junghun, et al.
Veröffentlicht: (2025)
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024) -
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023) -
Switching-Geometry Analysis of Deflated Q-Value Iteration
von: Lee, Donghwan
Veröffentlicht: (2026) -
Lyapunov-Certified Direct Switching Theory for Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026) -
A Discrete-Time Switching System Analysis of Q-learning
von: Lee, Donghwan, et al.
Veröffentlicht: (2021)