Non-stochastic Bandits With Evolving Observations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bar-On, Yogev, Mansour, Yishay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Regret for Policy Optimization in Contextual Bandits
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
Collaborating in Multi-Armed Bandits with Strategic Agents
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
von: Levy, Orin, et al.
Veröffentlicht: (2025)
von: Levy, Orin, et al.
Veröffentlicht: (2025)
The Real Price of Bandit Information in Multiclass Classification
von: Erez, Liad, et al.
Veröffentlicht: (2024)
von: Erez, Liad, et al.
Veröffentlicht: (2024)
Fast Rates for Bandit PAC Multiclass Classification
von: Erez, Liad, et al.
Veröffentlicht: (2024)
von: Erez, Liad, et al.
Veröffentlicht: (2024)
Rising Rested MAB with Linear Drift
von: Amichay, Omer, et al.
Veröffentlicht: (2025)
von: Amichay, Omer, et al.
Veröffentlicht: (2025)
Competing Bandits: The Perils of Exploration Under Competition
von: Aridor, Guy, et al.
Veröffentlicht: (2020)
von: Aridor, Guy, et al.
Veröffentlicht: (2020)
Modeling Attrition in Recommender Systems with Departing Bandits
von: Ben-Porat, Omer, et al.
Veröffentlicht: (2022)
von: Ben-Porat, Omer, et al.
Veröffentlicht: (2022)
How to Boost Any Loss Function
von: Nock, Richard, et al.
Veröffentlicht: (2024)
von: Nock, Richard, et al.
Veröffentlicht: (2024)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
von: Erez, Liad, et al.
Veröffentlicht: (2026)
von: Erez, Liad, et al.
Veröffentlicht: (2026)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
von: Dann, Christoph, et al.
Veröffentlicht: (2026)
von: Dann, Christoph, et al.
Veröffentlicht: (2026)
A Characterization of Semi-Supervised Adversarially-Robust PAC Learnability
von: Attias, Idan, et al.
Veröffentlicht: (2022)
von: Attias, Idan, et al.
Veröffentlicht: (2022)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
Delay as Payoff in MAB
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2024)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2024)
Probably Approximately Precision and Recall Learning
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Eluder-based Regret for Stochastic Contextual MDPs
von: Levy, Orin, et al.
Veröffentlicht: (2022)
von: Levy, Orin, et al.
Veröffentlicht: (2022)
The Hidden Cost of Approximation in Online Mirror Descent
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
Online Set Learning from Precision and Recall Feedback
von: Cohen, Lee, et al.
Veröffentlicht: (2026)
von: Cohen, Lee, et al.
Veröffentlicht: (2026)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
Bayesian Perspective on Memorization and Reconstruction
von: Kaplan, Haim, et al.
Veröffentlicht: (2025)
von: Kaplan, Haim, et al.
Veröffentlicht: (2025)
A Theoretical Framework for Statistical Evaluability of Generative Models
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
Learning-Augmented Algorithms with Explicit Predictors
von: Elias, Marek, et al.
Veröffentlicht: (2024)
von: Elias, Marek, et al.
Veröffentlicht: (2024)
Rate-Preserving Reductions for Blackwell Approachability
von: Dann, Christoph, et al.
Veröffentlicht: (2024)
von: Dann, Christoph, et al.
Veröffentlicht: (2024)
Cost-Aware Learning
von: Mohri, Clara, et al.
Veröffentlicht: (2026)
von: Mohri, Clara, et al.
Veröffentlicht: (2026)
Fast Inference via Hierarchical Speculative Decoding
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
Learnability Gaps of Strategic Classification
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
von: Chase, Zachary, et al.
Veröffentlicht: (2025)
von: Chase, Zachary, et al.
Veröffentlicht: (2025)
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms
von: Khosravi, Khashayar, et al.
Veröffentlicht: (2023)
von: Khosravi, Khashayar, et al.
Veröffentlicht: (2023)
Learning from Equivalence Queries, Revisited
von: Braverman, Mark, et al.
Veröffentlicht: (2026)
von: Braverman, Mark, et al.
Veröffentlicht: (2026)
On Mitigating Affinity Bias through Bandits with Evolving Biased Feedback
von: Faw, Matthew, et al.
Veröffentlicht: (2025)
von: Faw, Matthew, et al.
Veröffentlicht: (2025)
Sample Complexity of Agnostic Multiclass Classification: Natarajan Dimension Strikes Back
von: Cohen, Alon, et al.
Veröffentlicht: (2025)
von: Cohen, Alon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimal Regret for Policy Optimization in Contextual Bandits
von: Levy, Orin, et al.
Veröffentlicht: (2026) -
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025) -
Collaborating in Multi-Armed Bandits with Strategic Agents
von: Barnea, Idan, et al.
Veröffentlicht: (2026) -
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
von: Cassel, Asaf, et al.
Veröffentlicht: (2024) -
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
von: Barnea, Idan, et al.
Veröffentlicht: (2024)