Salvato in:
| Autori principali: | Yang, Hyukjun, Lim, Han-Dong, Lee, Donghwan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.06837 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence
di: Lee, Donghwan, et al.
Pubblicazione: (2026)
di: Lee, Donghwan, et al.
Pubblicazione: (2026)
Periodic Regularized Q-Learning
di: Yang, Hyukjun, et al.
Pubblicazione: (2026)
di: Yang, Hyukjun, et al.
Pubblicazione: (2026)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
A Switching System Theory of Q-Learning with Linear Function Approximation
di: Lee, Donghwan, et al.
Pubblicazione: (2026)
di: Lee, Donghwan, et al.
Pubblicazione: (2026)
Regularized Q-learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2022)
di: Lim, Han-Dong, et al.
Pubblicazione: (2022)
Backstepping Temporal Difference Learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
A primal-dual perspective for distributed TD-learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
A finite time analysis of distributed Q-learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
Analysis of approximate linear programming solution to Markov decision problem with log barrier function
di: Lee, Donghwan, et al.
Pubblicazione: (2025)
di: Lee, Donghwan, et al.
Pubblicazione: (2025)
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
Finite-Time Error Analysis of Soft Q-Learning: Switching System Approach
di: Jeong, Narim, et al.
Pubblicazione: (2024)
di: Jeong, Narim, et al.
Pubblicazione: (2024)
Distributional Off-policy Evaluation with Bellman Residual Minimization
di: Hong, Sungee, et al.
Pubblicazione: (2024)
di: Hong, Sungee, et al.
Pubblicazione: (2024)
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
New Versions of Gradient Temporal Difference Learning
di: Lee, Donghwan, et al.
Pubblicazione: (2021)
di: Lee, Donghwan, et al.
Pubblicazione: (2021)
Soft Deterministic Policy Gradient with Gaussian Smoothing
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
Stability and Generalization for Bellman Residuals
di: Kang, Enoch H., et al.
Pubblicazione: (2025)
di: Kang, Enoch H., et al.
Pubblicazione: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
di: Kumar, Navdeep, et al.
Pubblicazione: (2025)
di: Kumar, Navdeep, et al.
Pubblicazione: (2025)
Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation
di: Lee, Donghwan
Pubblicazione: (2024)
di: Lee, Donghwan
Pubblicazione: (2024)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
Safe-Support Q-Learning: Learning without Unsafe Exploration
di: Lim, Yeeun, et al.
Pubblicazione: (2026)
di: Lim, Yeeun, et al.
Pubblicazione: (2026)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
Finite-Time Analysis of Simultaneous Double Q-learning
di: Na, Hyunjun, et al.
Pubblicazione: (2024)
di: Na, Hyunjun, et al.
Pubblicazione: (2024)
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
di: Jeong, Narim, et al.
Pubblicazione: (2026)
di: Jeong, Narim, et al.
Pubblicazione: (2026)
Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes
di: Lee, Donghwan, et al.
Pubblicazione: (2023)
di: Lee, Donghwan, et al.
Pubblicazione: (2023)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
Merge and Bound: Direct Manipulations on Weights for Class Incremental Learning
di: Kim, Taehoon, et al.
Pubblicazione: (2025)
di: Kim, Taehoon, et al.
Pubblicazione: (2025)
Lyapunov-Certified Direct Switching Theory for Q-Learning
di: Lee, Donghwan
Pubblicazione: (2026)
di: Lee, Donghwan
Pubblicazione: (2026)
A Priori Sampling of Transition States with Guided Diffusion
di: Lim, Hyukjun, et al.
Pubblicazione: (2026)
di: Lim, Hyukjun, et al.
Pubblicazione: (2026)
Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
di: Lee, Taeho, et al.
Pubblicazione: (2026)
di: Lee, Taeho, et al.
Pubblicazione: (2026)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
di: Kohler, Hector, et al.
Pubblicazione: (2023)
di: Kohler, Hector, et al.
Pubblicazione: (2023)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
di: van der Laan, Lars, et al.
Pubblicazione: (2025)
di: van der Laan, Lars, et al.
Pubblicazione: (2025)
Adaptive Policy Backbone via Shared Network
di: Park, Bumgeun, et al.
Pubblicazione: (2025)
di: Park, Bumgeun, et al.
Pubblicazione: (2025)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
di: Lee, HyeAnn, et al.
Pubblicazione: (2023)
di: Lee, HyeAnn, et al.
Pubblicazione: (2023)
OTSS: Output-Targeted Soft Segmentation for Contextual Decision-Weight Learning
di: Hu, Renjun, et al.
Pubblicazione: (2026)
di: Hu, Renjun, et al.
Pubblicazione: (2026)
Bellman Error Centering
di: Chen, Xingguo, et al.
Pubblicazione: (2025)
di: Chen, Xingguo, et al.
Pubblicazione: (2025)
Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions
di: Lee, Donghwan, et al.
Pubblicazione: (2022)
di: Lee, Donghwan, et al.
Pubblicazione: (2022)
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
di: Rho, Donghwan
Pubblicazione: (2025)
di: Rho, Donghwan
Pubblicazione: (2025)
Documenti analoghi
-
Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence
di: Lee, Donghwan, et al.
Pubblicazione: (2026) -
Periodic Regularized Q-Learning
di: Yang, Hyukjun, et al.
Pubblicazione: (2026) -
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
di: Lim, Han-Dong, et al.
Pubblicazione: (2025) -
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
di: Lim, Han-Dong, et al.
Pubblicazione: (2023) -
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)