Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Donghwan, Na, Hyunjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Lyapunov-Certified Direct Switching Theory for Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Finite-Time Analysis of Simultaneous Double Q-learning
von: Na, Hyunjun, et al.
Veröffentlicht: (2024)
von: Na, Hyunjun, et al.
Veröffentlicht: (2024)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
von: Lee, HyeAnn, et al.
Veröffentlicht: (2023)
von: Lee, HyeAnn, et al.
Veröffentlicht: (2023)
Periodic Regularized Q-Learning
von: Yang, Hyukjun, et al.
Veröffentlicht: (2026)
von: Yang, Hyukjun, et al.
Veröffentlicht: (2026)
Safe-Support Q-Learning: Learning without Unsafe Exploration
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
von: Park, Jongchan, et al.
Veröffentlicht: (2025)
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
A finite time analysis of distributed Q-learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
Backstepping Temporal Difference Learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
von: Lee, Taeho, et al.
Veröffentlicht: (2026)
von: Lee, Taeho, et al.
Veröffentlicht: (2026)
On Tuning Neural ODE for Stability, Consistency and Faster Convergence
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2023)
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2023)
Adaptive Policy Backbone via Shared Network
von: Park, Bumgeun, et al.
Veröffentlicht: (2025)
von: Park, Bumgeun, et al.
Veröffentlicht: (2025)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Sign-Separated Finite-Time Error Analysis of Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
Certified Adversarial Robustness via Partition-based Randomized Smoothing
von: Goli, Hossein, et al.
Veröffentlicht: (2024)
von: Goli, Hossein, et al.
Veröffentlicht: (2024)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
von: Lee, Jeong Woon, et al.
Veröffentlicht: (2026)
von: Lee, Jeong Woon, et al.
Veröffentlicht: (2026)
HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors
von: Kim, Hyunjun
Veröffentlicht: (2025)
von: Kim, Hyunjun
Veröffentlicht: (2025)
Certified Robustness for Deep Equilibrium Models via Serialized Random Smoothing
von: Gao, Weizhi, et al.
Veröffentlicht: (2024)
von: Gao, Weizhi, et al.
Veröffentlicht: (2024)
Reconcile Certified Robustness and Accuracy for DNN-based Smoothed Majority Vote Classifier
von: Jin, Gaojie, et al.
Veröffentlicht: (2025)
von: Jin, Gaojie, et al.
Veröffentlicht: (2025)
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
von: Wang, Zixia, et al.
Veröffentlicht: (2025)
von: Wang, Zixia, et al.
Veröffentlicht: (2025)
Certified Training with Branch-and-Bound for Lyapunov-stable Neural Control
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates
von: Wang, Chengxiao, et al.
Veröffentlicht: (2026)
von: Wang, Chengxiao, et al.
Veröffentlicht: (2026)
Power Interpretable Causal ODE Networks: A Unified Model for Explainable Anomaly Detection and Root Cause Analysis in Power Systems
von: Sun, Yue, et al.
Veröffentlicht: (2026)
von: Sun, Yue, et al.
Veröffentlicht: (2026)
MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
Mitigating the Likelihood Paradox in Flow-based OOD Detection via Entropy Manipulation
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
von: He, Bowei, et al.
Veröffentlicht: (2025)
von: He, Bowei, et al.
Veröffentlicht: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
von: Lee, Vint, et al.
Veröffentlicht: (2023)
von: Lee, Vint, et al.
Veröffentlicht: (2023)
An Analysis under a Unified Fomulation of Learning Algorithms with Output Constraints
von: Song, Mooho, et al.
Veröffentlicht: (2024)
von: Song, Mooho, et al.
Veröffentlicht: (2024)
Frictional Q-Learning
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
Switching-Geometry Analysis of Deflated Q-Value Iteration
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
von: Han, Seungyub, et al.
Veröffentlicht: (2026)
von: Han, Seungyub, et al.
Veröffentlicht: (2026)
Chunk-Guided Q-Learning
von: Song, Gwanwoo, et al.
Veröffentlicht: (2026)
von: Song, Gwanwoo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026) -
Lyapunov-Certified Direct Switching Theory for Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026) -
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
von: Na, Hyunjun, et al.
Veröffentlicht: (2026) -
Finite-Time Analysis of Simultaneous Double Q-learning
von: Na, Hyunjun, et al.
Veröffentlicht: (2024) -
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
von: Lee, HyeAnn, et al.
Veröffentlicht: (2023)