Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moradipari, Ahmadreza, Pedramfar, Mohammad, Zini, Modjtaba Shokrian, Aggarwal, Vaneet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
A Unified Approach for Maximizing Continuous DR-submodular Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023)
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
von: Saxena, Naman, et al.
Veröffentlicht: (2026)
von: Saxena, Naman, et al.
Veröffentlicht: (2026)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
Sample-Efficient Constrained Reinforcement Learning with General Parameterization
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
A Unified Framework for Analyzing Meta-algorithms in Online Convex Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2023)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
von: Bai, Qinbo, et al.
Veröffentlicht: (2023)
von: Bai, Qinbo, et al.
Veröffentlicht: (2023)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes
von: Ganguly, Bhargav, et al.
Veröffentlicht: (2023)
von: Ganguly, Bhargav, et al.
Veröffentlicht: (2023)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
von: Aggarwal, Vaneet, et al.
Veröffentlicht: (2024)
von: Aggarwal, Vaneet, et al.
Veröffentlicht: (2024)
Discrete State Diffusion Models: A Sample Complexity Perspective
von: Srikanth, Aadithya, et al.
Veröffentlicht: (2025)
von: Srikanth, Aadithya, et al.
Veröffentlicht: (2025)
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2021)
von: Bai, Qinbo, et al.
Veröffentlicht: (2021)
Decentralized Projection-free Online Upper-Linearizable Optimization with Applications to DR-Submodular Optimization
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
von: Jain, Vineet, et al.
Veröffentlicht: (2025)
von: Jain, Vineet, et al.
Veröffentlicht: (2025)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2022)
von: Bai, Qinbo, et al.
Veröffentlicht: (2022)
Order-Optimal Sample Complexity of Rectified Flows
von: Sahoo, Hari Krishna, et al.
Veröffentlicht: (2026)
von: Sahoo, Hari Krishna, et al.
Veröffentlicht: (2026)
Kernelized Reinforcement Learning with Order Optimal Regret Bounds
von: Vakili, Sattar, et al.
Veröffentlicht: (2023)
von: Vakili, Sattar, et al.
Veröffentlicht: (2023)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
On Regret Bounds of Thompson Sampling for Bayesian Optimization
von: Takeno, Shion, et al.
Veröffentlicht: (2026)
von: Takeno, Shion, et al.
Veröffentlicht: (2026)
Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
von: Xu, Mengfan, et al.
Veröffentlicht: (2020)
von: Xu, Mengfan, et al.
Veröffentlicht: (2020)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
von: Gaur, Mudit, et al.
Veröffentlicht: (2024)
Variational Offline Multi-agent Skill Discovery
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
Bridging Distributional and Risk-sensitive Reinforcement Learning with Provable Regret Bounds
von: Liang, Hao, et al.
Veröffentlicht: (2022)
von: Liang, Hao, et al.
Veröffentlicht: (2022)
Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
von: Zhang, Jiefu, et al.
Veröffentlicht: (2026)
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
von: Gaur, Mudit, et al.
Veröffentlicht: (2025)
Augmenting generative models with biomedical knowledge graphs improves targeted drug discovery
von: Malusare, Aditya, et al.
Veröffentlicht: (2025)
von: Malusare, Aditya, et al.
Veröffentlicht: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Reinforced Sequential Decision-Making for Sepsis Treatment: The POSNEGDM Framework with Mortality Classifier and Transformer
von: Tamboli, Dipesh, et al.
Veröffentlicht: (2024)
von: Tamboli, Dipesh, et al.
Veröffentlicht: (2024)
Graph Neural Thompson Sampling
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
von: Fourati, Fares, et al.
Veröffentlicht: (2025)
von: Fourati, Fares, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026) -
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023) -
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
von: Lu, Yiyang, et al.
Veröffentlicht: (2025) -
A Unified Approach for Maximizing Continuous DR-submodular Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2023) -
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
von: Saxena, Naman, et al.
Veröffentlicht: (2026)