Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Levy, Orin, Erez, Liad, Cohen, Alon, Mansour, Yishay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Eluder-based Regret for Stochastic Contextual MDPs
by: Levy, Orin, et al.
Published: (2022)
by: Levy, Orin, et al.
Published: (2022)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
The Real Price of Bandit Information in Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
by: Erez, Liad, et al.
Published: (2022)
by: Erez, Liad, et al.
Published: (2022)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
by: Erez, Liad, et al.
Published: (2025)
by: Erez, Liad, et al.
Published: (2025)
Sample Complexity of Agnostic Multiclass Classification: Natarajan Dimension Strikes Back
by: Cohen, Alon, et al.
Published: (2025)
by: Cohen, Alon, et al.
Published: (2025)
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
by: Barnea, Idan, et al.
Published: (2024)
by: Barnea, Idan, et al.
Published: (2024)
Non-stochastic Bandits With Evolving Observations
by: Bar-On, Yogev, et al.
Published: (2024)
by: Bar-On, Yogev, et al.
Published: (2024)
Regret Guarantees for Linear Contextual Stochastic Shortest Path
by: Polikar, Dor, et al.
Published: (2025)
by: Polikar, Dor, et al.
Published: (2025)
Delay as Payoff in MAB
by: Schlisselberg, Ofir, et al.
Published: (2024)
by: Schlisselberg, Ofir, et al.
Published: (2024)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
by: Sherman, Uri, et al.
Published: (2023)
by: Sherman, Uri, et al.
Published: (2023)
Improved Regret for Bandit Convex Optimization with Delayed Feedback
by: Wan, Yuanyu, et al.
Published: (2024)
by: Wan, Yuanyu, et al.
Published: (2024)
Probably Approximately Precision and Recall Learning
by: Cohen, Lee, et al.
Published: (2024)
by: Cohen, Lee, et al.
Published: (2024)
Online Set Learning from Precision and Recall Feedback
by: Cohen, Lee, et al.
Published: (2026)
by: Cohen, Lee, et al.
Published: (2026)
Distributed Learning in Markovian Restless Bandits over Interference Graphs for Stable Spectrum Sharing
by: Didi, Liad Lea, et al.
Published: (2025)
by: Didi, Liad Lea, et al.
Published: (2025)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Collaborating in Multi-Armed Bandits with Strategic Agents
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Information Capacity Regret Bounds for Bandits with Mediator Feedback
by: Eldowa, Khaled, et al.
Published: (2024)
by: Eldowa, Khaled, et al.
Published: (2024)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Second Order Bounds for Contextual Bandits with Function Approximation
by: Pacchiano, Aldo
Published: (2024)
by: Pacchiano, Aldo
Published: (2024)
Neural Contextual Bandits Under Delayed Feedback Constraints
by: Moghimi, Mohammadali, et al.
Published: (2025)
by: Moghimi, Mohammadali, et al.
Published: (2025)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Modeling Attrition in Recommender Systems with Departing Bandits
by: Ben-Porat, Omer, et al.
Published: (2022)
by: Ben-Porat, Omer, et al.
Published: (2022)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
A Characterization of Semi-Supervised Adversarially-Robust PAC Learnability
by: Attias, Idan, et al.
Published: (2022)
by: Attias, Idan, et al.
Published: (2022)
How to Boost Any Loss Function
by: Nock, Richard, et al.
Published: (2024)
by: Nock, Richard, et al.
Published: (2024)
Fast Best-in-Class Regret for Contextual Bandits
by: Girard, Samuel, et al.
Published: (2025)
by: Girard, Samuel, et al.
Published: (2025)
Adversarial Bandits with Multi-User Delayed Feedback: Theory and Application
by: Li, Yandi, et al.
Published: (2023)
by: Li, Yandi, et al.
Published: (2023)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
The Hidden Cost of Approximation in Online Mirror Descent
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback
by: Huang, Ruiyuan, et al.
Published: (2025)
by: Huang, Ruiyuan, et al.
Published: (2025)
How Does Variance Shape the Regret in Contextual Bandits?
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
Similar Items
-
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026) -
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026) -
Eluder-based Regret for Stochastic Contextual MDPs
by: Levy, Orin, et al.
Published: (2022) -
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024) -
The Real Price of Bandit Information in Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)