Bayesian Regret Minimization in Offline Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Petrik, Marek, Tennenholtz, Guy, Ghavamzadeh, Mohammad |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
by: Hau, Jia Lin, et al.
Published: (2022)
by: Hau, Jia Lin, et al.
Published: (2022)
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
by: Hau, Jia Lin, et al.
Published: (2024)
by: Hau, Jia Lin, et al.
Published: (2024)
Contextual Bandits with Stage-wise Constraints
by: Pacchiano, Aldo, et al.
Published: (2024)
by: Pacchiano, Aldo, et al.
Published: (2024)
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators
by: Linial, Ori, et al.
Published: (2024)
by: Linial, Ori, et al.
Published: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024)
by: Deb, Rohan, et al.
Published: (2024)
Bayesian Robust Optimization for Imitation Learning
by: Brown, Daniel S., et al.
Published: (2020)
by: Brown, Daniel S., et al.
Published: (2020)
Efficient Swap Regret Minimization in Combinatorial Bandits
by: Kontogiannis, Andreas, et al.
Published: (2026)
by: Kontogiannis, Andreas, et al.
Published: (2026)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
by: Bernasconi, Martino, et al.
Published: (2024)
by: Bernasconi, Martino, et al.
Published: (2024)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Percentile Criterion Optimization in Offline Reinforcement Learning
by: Lobo, Elita A., et al.
Published: (2024)
by: Lobo, Elita A., et al.
Published: (2024)
Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution
by: Feng, Qing, et al.
Published: (2024)
by: Feng, Qing, et al.
Published: (2024)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
by: Grand-Clément, Julien, et al.
Published: (2023)
by: Grand-Clément, Julien, et al.
Published: (2023)
Graph-Dependent Regret Bounds in Multi-Armed Bandits with Interference
by: Jamshidi, Fateme, et al.
Published: (2025)
by: Jamshidi, Fateme, et al.
Published: (2025)
Active Context Selection Improves Simple Regret in Contextual Bandits
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
by: Hou, Yunlong, et al.
Published: (2026)
by: Hou, Yunlong, et al.
Published: (2026)
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2025)
by: Genalti, Gianmarco, et al.
Published: (2025)
Optimal Regret for Single Index Bandits
by: Dey, Devdan, et al.
Published: (2026)
by: Dey, Devdan, et al.
Published: (2026)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
by: Xue, Bo, et al.
Published: (2025)
by: Xue, Bo, et al.
Published: (2025)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Representation-Driven Reinforcement Learning
by: Nabati, Ofir, et al.
Published: (2023)
by: Nabati, Ofir, et al.
Published: (2023)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Fast Best-in-Class Regret for Contextual Bandits
by: Girard, Samuel, et al.
Published: (2025)
by: Girard, Samuel, et al.
Published: (2025)
Improved Regret Bounds for Bandits with Expert Advice
by: Cesa-Bianchi, Nicolò, et al.
Published: (2024)
by: Cesa-Bianchi, Nicolò, et al.
Published: (2024)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Unlearning Offline Stochastic Multi-Armed Bandits
by: Ye, Zichun, et al.
Published: (2026)
by: Ye, Zichun, et al.
Published: (2026)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
Spectral Bellman Method: Unifying Representation and Exploration in RL
by: Nabati, Ofir, et al.
Published: (2025)
by: Nabati, Ofir, et al.
Published: (2025)
Adjusted Expected Improvement for Cumulative Regret Minimization in Noisy Bayesian Optimization
by: Hu, Shouri, et al.
Published: (2022)
by: Hu, Shouri, et al.
Published: (2022)
Group-Sensitive Offline Contextual Bandits
by: Guo, Yihong, et al.
Published: (2025)
by: Guo, Yihong, et al.
Published: (2025)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Parameter-Free Dynamic Regret for Unconstrained Linear Bandits
by: Rumi, Alberto, et al.
Published: (2026)
by: Rumi, Alberto, et al.
Published: (2026)
How Does Variance Shape the Regret in Contextual Bandits?
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
Improved Regret for Bandit Convex Optimization with Delayed Feedback
by: Wan, Yuanyu, et al.
Published: (2024)
by: Wan, Yuanyu, et al.
Published: (2024)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
by: Shufaro, Itai, et al.
Published: (2024)
by: Shufaro, Itai, et al.
Published: (2024)
No-Regret Linear Bandits under Gap-Adjusted Misspecification
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
Similar Items
-
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
by: Hau, Jia Lin, et al.
Published: (2022) -
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
by: Hau, Jia Lin, et al.
Published: (2024) -
Contextual Bandits with Stage-wise Constraints
by: Pacchiano, Aldo, et al.
Published: (2024) -
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators
by: Linial, Ori, et al.
Published: (2024) -
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024)