Soft Deterministic Policy Gradient with Gaussian Smoothing
Fuente:
arXiv
Saved in:
| Main Authors: | Na, Hyunjun, Lee, Donghwan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
by: Lee, Donghwan, et al.
Published: (2024)
by: Lee, Donghwan, et al.
Published: (2024)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
by: Lee, Taeho, et al.
Published: (2026)
by: Lee, Taeho, et al.
Published: (2026)
Adaptive Policy Backbone via Shared Network
by: Park, Bumgeun, et al.
Published: (2025)
by: Park, Bumgeun, et al.
Published: (2025)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Finite-Time Analysis of Simultaneous Double Q-learning
by: Na, Hyunjun, et al.
Published: (2024)
by: Na, Hyunjun, et al.
Published: (2024)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
Robust Deterministic Policy Gradient for Disturbance Attenuation and Its Application to Quadrotor Control
by: Lee, Taeho, et al.
Published: (2025)
by: Lee, Taeho, et al.
Published: (2025)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
Lyapunov-Certified Direct Switching Theory for Q-Learning
by: Lee, Donghwan
Published: (2026)
by: Lee, Donghwan
Published: (2026)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Suppressing Overestimation in Q-Learning through Adversarial Behaviors
by: Lee, HyeAnn, et al.
Published: (2023)
by: Lee, HyeAnn, et al.
Published: (2023)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Edge Delayed Deep Deterministic Policy Gradient: efficient continuous control for edge scenarios
by: Sinigaglia, Alberto, et al.
Published: (2024)
by: Sinigaglia, Alberto, et al.
Published: (2024)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Backstepping Temporal Difference Learning
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors
by: Kim, Hyunjun
Published: (2025)
by: Kim, Hyunjun
Published: (2025)
Safe-Support Q-Learning: Learning without Unsafe Exploration
by: Lim, Yeeun, et al.
Published: (2026)
by: Lim, Yeeun, et al.
Published: (2026)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
by: Park, Jongchan, et al.
Published: (2025)
by: Park, Jongchan, et al.
Published: (2025)
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
by: Lim, Han-Dong, et al.
Published: (2024)
by: Lim, Han-Dong, et al.
Published: (2024)
Axiomatization of Gradient Smoothing in Neural Networks
by: Zhou, Linjiang, et al.
Published: (2024)
by: Zhou, Linjiang, et al.
Published: (2024)
Periodic Regularized Q-Learning
by: Yang, Hyukjun, et al.
Published: (2026)
by: Yang, Hyukjun, et al.
Published: (2026)
MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse
by: Kim, Donghwan, et al.
Published: (2026)
by: Kim, Donghwan, et al.
Published: (2026)
Mitigating the Likelihood Paradox in Flow-based OOD Detection via Entropy Manipulation
by: Kim, Donghwan, et al.
Published: (2026)
by: Kim, Donghwan, et al.
Published: (2026)
A finite time analysis of distributed Q-learning
by: Lim, Han-Dong, et al.
Published: (2024)
by: Lim, Han-Dong, et al.
Published: (2024)
Direct Soft-Policy Sampling via Langevin Dynamics
by: Ki, Donghyeon, et al.
Published: (2026)
by: Ki, Donghyeon, et al.
Published: (2026)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Soft Sequence Policy Optimization
by: Glazyrina, Svetlana, et al.
Published: (2026)
by: Glazyrina, Svetlana, et al.
Published: (2026)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
by: Lee, Hosu, et al.
Published: (2024)
by: Lee, Hosu, et al.
Published: (2024)
Epidemic Control on a Large-Scale-Agent-Based Epidemiology Model using Deep Deterministic Policy Gradient
by: Deshkar, Gaurav, et al.
Published: (2023)
by: Deshkar, Gaurav, et al.
Published: (2023)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Learning Generalized Policies for Fully Observable Non-Deterministic Planning Domains
by: Hofmann, Till, et al.
Published: (2024)
by: Hofmann, Till, et al.
Published: (2024)
Why the Counterintuitive Phenomenon of Likelihood Rarely Appears in Tabular Anomaly Detection with Deep Generative Models?
by: Kim, Donghwan, et al.
Published: (2026)
by: Kim, Donghwan, et al.
Published: (2026)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
by: Vo, Thanh Vinh, et al.
Published: (2025)
by: Vo, Thanh Vinh, et al.
Published: (2025)
Federated Smoothing Proximal Gradient for Quantile Regression with Non-Convex Penalties
by: Mirzaeifard, Reza, et al.
Published: (2024)
by: Mirzaeifard, Reza, et al.
Published: (2024)
Similar Items
-
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
by: Lee, Donghwan, et al.
Published: (2024) -
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
by: Na, Hyunjun, et al.
Published: (2026) -
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
by: Lee, Taeho, et al.
Published: (2026) -
Adaptive Policy Backbone via Shared Network
by: Park, Bumgeun, et al.
Published: (2025) -
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)