Efficient Exploration in Average-Reward Constrained Reinforcement Learning: Achieving Near-Optimal Regret With Posterior Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Provodin, Danil, Kaptein, Maurits, Pechenizkiy, Mykola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Knowledge Transfer in Learning Using Privileged Information
by: Provodin, Danil, et al.
Published: (2024)
by: Provodin, Danil, et al.
Published: (2024)
A Verifier Hierarchy
by: Kaptein, Maurits
Published: (2025)
by: Kaptein, Maurits
Published: (2025)
Incorporating structural uncertainty in causal decision making
by: Kaptein, Maurits
Published: (2025)
by: Kaptein, Maurits
Published: (2025)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Logarithmic Regret of Exploration in Average Reward Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025)
by: Boone, Victor, et al.
Published: (2025)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
by: Roch, Zachary, et al.
Published: (2025)
by: Roch, Zachary, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Bandits for Sponsored Search Auctions under Unknown Valuation Model: Case Study in E-Commerce Advertising
by: Provodin, Danil, et al.
Published: (2023)
by: Provodin, Danil, et al.
Published: (2023)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
by: Tomilin, Tristan, et al.
Published: (2025)
by: Tomilin, Tristan, et al.
Published: (2025)
Achieving $\widetilde{\mathcal{O}}(\sqrt{T})$ Regret in Average-Reward POMDPs with Known Observation Models
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
Near-Optimal Sample Complexity in Reward-Free Kernel-Based Reinforcement Learning
by: Kayal, Aya, et al.
Published: (2025)
by: Kayal, Aya, et al.
Published: (2025)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024)
by: Vakili, Sattar, et al.
Published: (2024)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Online Generalized-mean Welfare Maximization: Achieving Near-Optimal Regret from Samples
by: Yang, Zongjun, et al.
Published: (2026)
by: Yang, Zongjun, et al.
Published: (2026)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
Are There Exceptions to Goodhart's Law? On the Moral Justification of Fairness-Aware Machine Learning
by: Weerts, Hilde, et al.
Published: (2022)
by: Weerts, Hilde, et al.
Published: (2022)
Adaptive Sparsity Level during Training for Efficient Time Series Forecasting with Transformers
by: Atashgahi, Zahra, et al.
Published: (2023)
by: Atashgahi, Zahra, et al.
Published: (2023)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
by: Flynn, Hamish, et al.
Published: (2026)
by: Flynn, Hamish, et al.
Published: (2026)
Robust Active Learning (RoAL): Countering Dynamic Adversaries in Active Learning with Elastic Weight Consolidation
by: Fajri, Ricky Maulana, et al.
Published: (2024)
by: Fajri, Ricky Maulana, et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning
by: Jiao, Yuchen, et al.
Published: (2026)
by: Jiao, Yuchen, et al.
Published: (2026)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
by: Aggarwal, Vaneet, et al.
Published: (2024)
by: Aggarwal, Vaneet, et al.
Published: (2024)
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
One-Shot Federated Learning with Bayesian Pseudocoresets
by: d'Hondt, Tim, et al.
Published: (2024)
by: d'Hondt, Tim, et al.
Published: (2024)
Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Near-Optimal Sample Complexity for Online Constrained MDPs
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Boosting Robustness in Preference-Based Reinforcement Learning with Dynamic Sparsity
by: Muslimani, Calarina, et al.
Published: (2024)
by: Muslimani, Calarina, et al.
Published: (2024)
Conformalized Exceptional Model Mining: Telling Where Your Model Performs (Not) Well
by: Du, Xin, et al.
Published: (2025)
by: Du, Xin, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
Gaussian Process Upper Confidence Bound Achieves Nearly-Optimal Regret in Noise-Free Gaussian Process Bandits
by: Iwazaki, Shogo
Published: (2025)
by: Iwazaki, Shogo
Published: (2025)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling, and Beyond
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
Posterior Sampling-Based Bayesian Optimization with Tighter Bayesian Regret Bounds
by: Takeno, Shion, et al.
Published: (2023)
by: Takeno, Shion, et al.
Published: (2023)
Regret Analysis of Posterior Sampling-Based Expected Improvement for Bayesian Optimization
by: Takeno, Shion, et al.
Published: (2025)
by: Takeno, Shion, et al.
Published: (2025)
Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits
by: Ye, Zichun, et al.
Published: (2025)
by: Ye, Zichun, et al.
Published: (2025)
Similar Items
-
Rethinking Knowledge Transfer in Learning Using Privileged Information
by: Provodin, Danil, et al.
Published: (2024) -
A Verifier Hierarchy
by: Kaptein, Maurits
Published: (2025) -
Incorporating structural uncertainty in causal decision making
by: Kaptein, Maurits
Published: (2025) -
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025) -
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)