Optimal Regret for Policy Optimization in Contextual Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Levy, Orin, Mansour, Yishay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
by: Levy, Orin, et al.
Published: (2025)
by: Levy, Orin, et al.
Published: (2025)
Eluder-based Regret for Stochastic Contextual MDPs
by: Levy, Orin, et al.
Published: (2022)
by: Levy, Orin, et al.
Published: (2022)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
by: Barnea, Idan, et al.
Published: (2024)
by: Barnea, Idan, et al.
Published: (2024)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
Non-stochastic Bandits With Evolving Observations
by: Bar-On, Yogev, et al.
Published: (2024)
by: Bar-On, Yogev, et al.
Published: (2024)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
by: Sherman, Uri, et al.
Published: (2023)
by: Sherman, Uri, et al.
Published: (2023)
Collaborating in Multi-Armed Bandits with Strategic Agents
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
On the Optimal Regret of Locally Private Linear Contextual Bandit
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
by: Dann, Christoph, et al.
Published: (2026)
by: Dann, Christoph, et al.
Published: (2026)
Fast Best-in-Class Regret for Contextual Bandits
by: Girard, Samuel, et al.
Published: (2025)
by: Girard, Samuel, et al.
Published: (2025)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
How Does Variance Shape the Regret in Contextual Bandits?
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
The Real Price of Bandit Information in Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Rising Rested MAB with Linear Drift
by: Amichay, Omer, et al.
Published: (2025)
by: Amichay, Omer, et al.
Published: (2025)
Optimal Regret for Single Index Bandits
by: Dey, Devdan, et al.
Published: (2026)
by: Dey, Devdan, et al.
Published: (2026)
Active Context Selection Improves Simple Regret in Contextual Bandits
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Competing Bandits: The Perils of Exploration Under Competition
by: Aridor, Guy, et al.
Published: (2020)
by: Aridor, Guy, et al.
Published: (2020)
Optimal Baseline Corrections for Off-Policy Contextual Bandits
by: Gupta, Shashank, et al.
Published: (2024)
by: Gupta, Shashank, et al.
Published: (2024)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
by: Erez, Liad, et al.
Published: (2022)
by: Erez, Liad, et al.
Published: (2022)
Modeling Attrition in Recommender Systems with Departing Bandits
by: Ben-Porat, Omer, et al.
Published: (2022)
by: Ben-Porat, Omer, et al.
Published: (2022)
How to Boost Any Loss Function
by: Nock, Richard, et al.
Published: (2024)
by: Nock, Richard, et al.
Published: (2024)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
by: Kim, Seok-Jin, et al.
Published: (2024)
by: Kim, Seok-Jin, et al.
Published: (2024)
Optimal High-Probability Regret for Online Convex Optimization with Two-Point Bandit Feedback
by: Ye, Haishan
Published: (2026)
by: Ye, Haishan
Published: (2026)
Swap Regret and Correlated Equilibria Beyond Normal-Form Games
by: Arunachaleswaran, Eshwar Ram, et al.
Published: (2025)
by: Arunachaleswaran, Eshwar Ram, et al.
Published: (2025)
Online Weighted Paging with Unknown Weights
by: Levy, Orin, et al.
Published: (2024)
by: Levy, Orin, et al.
Published: (2024)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
A Characterization of Semi-Supervised Adversarially-Robust PAC Learnability
by: Attias, Idan, et al.
Published: (2022)
by: Attias, Idan, et al.
Published: (2022)
Preference-centric Bandits: Optimality of Mixtures and Regret-efficient Algorithms
by: Tatlı, Meltem, et al.
Published: (2025)
by: Tatlı, Meltem, et al.
Published: (2025)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Similar Items
-
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026) -
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024) -
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
by: Levy, Orin, et al.
Published: (2025) -
Eluder-based Regret for Stochastic Contextual MDPs
by: Levy, Orin, et al.
Published: (2022) -
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)