Suppressing Overestimation in Q-Learning through Adversarial Behaviors
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, HyeAnn, Lee, Donghwan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
Lyapunov-Certified Direct Switching Theory for Q-Learning
di: Lee, Donghwan
Pubblicazione: (2026)
di: Lee, Donghwan
Pubblicazione: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
di: Lee, Taeho, et al.
Pubblicazione: (2026)
di: Lee, Taeho, et al.
Pubblicazione: (2026)
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
Safe-Support Q-Learning: Learning without Unsafe Exploration
di: Lim, Yeeun, et al.
Pubblicazione: (2026)
di: Lim, Yeeun, et al.
Pubblicazione: (2026)
Periodic Regularized Q-Learning
di: Yang, Hyukjun, et al.
Pubblicazione: (2026)
di: Yang, Hyukjun, et al.
Pubblicazione: (2026)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
di: Park, Jongchan, et al.
Pubblicazione: (2025)
di: Park, Jongchan, et al.
Pubblicazione: (2025)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
A finite time analysis of distributed Q-learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
di: Lim, Han-Dong, et al.
Pubblicazione: (2024)
Adversarial Robustness Overestimation and Instability in TRADES
di: Li, Jonathan Weiping, et al.
Pubblicazione: (2024)
di: Li, Jonathan Weiping, et al.
Pubblicazione: (2024)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
Backstepping Temporal Difference Learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
di: Lim, Han-Dong, et al.
Pubblicazione: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
di: Na, Hyunjun, et al.
Pubblicazione: (2026)
Adaptive Policy Backbone via Shared Network
di: Park, Bumgeun, et al.
Pubblicazione: (2025)
di: Park, Bumgeun, et al.
Pubblicazione: (2025)
Sign-Separated Finite-Time Error Analysis of Q-Learning
di: Lee, Donghwan
Pubblicazione: (2026)
di: Lee, Donghwan
Pubblicazione: (2026)
QSIM: Mitigating Overestimation in Multi-Agent Reinforcement Learning via Action Similarity Weighted Q-Learning
di: Li, Yuanjun, et al.
Pubblicazione: (2026)
di: Li, Yuanjun, et al.
Pubblicazione: (2026)
Feature-Enhanced Machine Learning for All-Cause Mortality Prediction in Healthcare Data
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
Frictional Q-Learning
di: Kim, Hyunwoo, et al.
Pubblicazione: (2025)
di: Kim, Hyunwoo, et al.
Pubblicazione: (2025)
Chunk-Guided Q-Learning
di: Song, Gwanwoo, et al.
Pubblicazione: (2026)
di: Song, Gwanwoo, et al.
Pubblicazione: (2026)
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning
di: Park, Kwanyoung, et al.
Pubblicazione: (2024)
di: Park, Kwanyoung, et al.
Pubblicazione: (2024)
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
di: Lee, Dohyeok, et al.
Pubblicazione: (2024)
di: Lee, Dohyeok, et al.
Pubblicazione: (2024)
Sample-efficient Adversarial Imitation Learning
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
Learning Robust Reasoning through Guided Adversarial Self-Play
di: Li, Shuozhe, et al.
Pubblicazione: (2026)
di: Li, Shuozhe, et al.
Pubblicazione: (2026)
Switching-Geometry Analysis of Deflated Q-Value Iteration
di: Lee, Donghwan
Pubblicazione: (2026)
di: Lee, Donghwan
Pubblicazione: (2026)
Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
di: Park, Inkyu, et al.
Pubblicazione: (2025)
di: Park, Inkyu, et al.
Pubblicazione: (2025)
MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
Mitigating the Likelihood Paradox in Flow-based OOD Detection via Entropy Manipulation
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
di: Yeom, Junghyuk, et al.
Pubblicazione: (2024)
di: Yeom, Junghyuk, et al.
Pubblicazione: (2024)
$β$-DQN: Improving Deep Q-Learning By Evolving the Behavior
di: Zhang, Hongming, et al.
Pubblicazione: (2025)
di: Zhang, Hongming, et al.
Pubblicazione: (2025)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
di: Dalal, Gal, et al.
Pubblicazione: (2026)
di: Dalal, Gal, et al.
Pubblicazione: (2026)
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
di: Yamabe, Shojiro, et al.
Pubblicazione: (2024)
di: Yamabe, Shojiro, et al.
Pubblicazione: (2024)
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
di: Doo, JaeHyeok, et al.
Pubblicazione: (2026)
di: Doo, JaeHyeok, et al.
Pubblicazione: (2026)
DAFA: Distance-Aware Fair Adversarial Training
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
di: Lee, Hyungyu, et al.
Pubblicazione: (2024)
Finite-Time Error Analysis of Soft Q-Learning: Switching System Approach
di: Jeong, Narim, et al.
Pubblicazione: (2024)
di: Jeong, Narim, et al.
Pubblicazione: (2024)
CASA: CNN Autoencoder-based Score Attention for Efficient Multivariate Long-term Time-series Forecasting
di: Lee, Minhyuk, et al.
Pubblicazione: (2025)
di: Lee, Minhyuk, et al.
Pubblicazione: (2025)
Why the Counterintuitive Phenomenon of Likelihood Rarely Appears in Tabular Anomaly Detection with Deep Generative Models?
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
di: Kim, Donghwan, et al.
Pubblicazione: (2026)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
Stochastic Actor-Critic: Mitigating Overestimation via Temporal Aleatoric Uncertainty
di: Özalp, Uğurcan
Pubblicazione: (2026)
di: Özalp, Uğurcan
Pubblicazione: (2026)
Documenti analoghi
-
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
di: Lim, Han-Dong, et al.
Pubblicazione: (2024) -
Lyapunov-Certified Direct Switching Theory for Q-Learning
di: Lee, Donghwan
Pubblicazione: (2026) -
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
di: Lee, Taeho, et al.
Pubblicazione: (2026) -
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
di: Lee, Donghwan, et al.
Pubblicazione: (2024) -
Safe-Support Q-Learning: Learning without Unsafe Exploration
di: Lim, Yeeun, et al.
Pubblicazione: (2026)