Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Kihyun, Lee, Duksang, Overman, William, Lee, Dabeen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Parameter-Free Algorithms for Performative Regret Minimization under Decision-Dependent Distributions
by: Park, Sungwoo, et al.
Published: (2024)
by: Park, Sungwoo, et al.
Published: (2024)
Stochastic-Constrained Stochastic Optimization with Markovian Data
by: Kim, Yeongjong, et al.
Published: (2023)
by: Kim, Yeongjong, et al.
Published: (2023)
Reinforcement Learning and Regret Bounds for Admission Control
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
Optimism as Risk-Seeking in Multi-Agent Reinforcement Learning
by: Zhang, Runyu, et al.
Published: (2025)
by: Zhang, Runyu, et al.
Published: (2025)
Faster Rates for No-Regret Learning in General Games via Cautious Optimism
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
by: Zhang, Dake, et al.
Published: (2024)
by: Zhang, Dake, et al.
Published: (2024)
Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double Optimism
by: Patitucci, Francisco, et al.
Published: (2025)
by: Patitucci, Francisco, et al.
Published: (2025)
Online (Non-)Convex Learning via Tempered Optimism
by: Haddouche, Maxime, et al.
Published: (2023)
by: Haddouche, Maxime, et al.
Published: (2023)
Chebyshev Center-Based Direction Selection for Multi-Objective Optimization and Training PINNs
by: Yoon, Hoyeol, et al.
Published: (2026)
by: Yoon, Hoyeol, et al.
Published: (2026)
Tail Distribution of Regret in Optimistic Reinforcement Learning
by: Khodadadian, Sajad, et al.
Published: (2025)
by: Khodadadian, Sajad, et al.
Published: (2025)
Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems
by: Ziemann, Ingvar, et al.
Published: (2022)
by: Ziemann, Ingvar, et al.
Published: (2022)
Rate-Optimal Regret for the Safe Learning-based Control of the Constrained Linear Quadratic Regulator
by: Hutchinson, Spencer, et al.
Published: (2026)
by: Hutchinson, Spencer, et al.
Published: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Cautious Optimism: A Meta-Algorithm for Near-Constant Regret in General Games
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Tighter Performance Theory of FedExProx
by: Anyszka, Wojciech, et al.
Published: (2024)
by: Anyszka, Wojciech, et al.
Published: (2024)
Regret Bounds for Episodic Risk-Sensitive Linear Quadratic Regulator
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Implicit Riemannian Optimism with Applications to Min-Max Problems
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Regret Bounds for Expected Improvement Algorithms in Gaussian Process Bandit Optimization
by: Tran-The, Hung, et al.
Published: (2022)
by: Tran-The, Hung, et al.
Published: (2022)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
by: Zhou, Angela
Published: (2024)
by: Zhou, Angela
Published: (2024)
Achieving Better Local Regret Bound for Online Non-Convex Bilevel Optimization
by: Jia, Tingkai, et al.
Published: (2026)
by: Jia, Tingkai, et al.
Published: (2026)
Optimistic Online LQR via Intrinsic Rewards
by: Bartos, Marcell, et al.
Published: (2026)
by: Bartos, Marcell, et al.
Published: (2026)
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
by: Sattar, Yahya, et al.
Published: (2021)
by: Sattar, Yahya, et al.
Published: (2021)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
by: Ozdaglar, Asuman, et al.
Published: (2022)
by: Ozdaglar, Asuman, et al.
Published: (2022)
Sublinear Regret for a Class of Continuous-Time Linear-Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2024)
by: Huang, Yilie, et al.
Published: (2024)
Sampling-based Safe Reinforcement Learning for Nonlinear Dynamical Systems
by: Suttle, Wesley A., et al.
Published: (2024)
by: Suttle, Wesley A., et al.
Published: (2024)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
Online Optimization on Hadamard Manifolds: Curvature Independent Regret Bounds on Horospherically Convex Objectives
by: Sahinoglu, Emre, et al.
Published: (2025)
by: Sahinoglu, Emre, et al.
Published: (2025)
Tight Regret Bounds for Bayesian Optimization in One Dimension
by: Scarlett, Jonathan
Published: (2018)
by: Scarlett, Jonathan
Published: (2018)
Methodology for Interpretable Reinforcement Learning for Optimizing Mechanical Ventilation
by: Lee, Joo Seung, et al.
Published: (2024)
by: Lee, Joo Seung, et al.
Published: (2024)
Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time
by: Mazumdar, Abhijit, et al.
Published: (2024)
by: Mazumdar, Abhijit, et al.
Published: (2024)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
by: Xie, Yan-Feng, et al.
Published: (2026)
by: Xie, Yan-Feng, et al.
Published: (2026)
Learning Decentralized Linear Quadratic Regulators with $\sqrt{T}$ Regret
by: Ye, Lintao, et al.
Published: (2022)
by: Ye, Lintao, et al.
Published: (2022)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
by: Zamir, Guy, et al.
Published: (2026)
by: Zamir, Guy, et al.
Published: (2026)
Similar Items
-
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026) -
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
by: Yu, Kihyun, et al.
Published: (2026) -
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024) -
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024) -
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
by: Park, Jaehyun, et al.
Published: (2024)