Eluder-based Regret for Stochastic Contextual MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Levy, Orin, Cassel, Asaf, Cohen, Alon, Mansour, Yishay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Optimal Regret for Policy Optimization in Contextual Bandits
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
von: Levy, Orin, et al.
Veröffentlicht: (2025)
von: Levy, Orin, et al.
Veröffentlicht: (2025)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Regret Guarantees for Linear Contextual Stochastic Shortest Path
von: Polikar, Dor, et al.
Veröffentlicht: (2025)
von: Polikar, Dor, et al.
Veröffentlicht: (2025)
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
von: Erez, Liad, et al.
Veröffentlicht: (2026)
von: Erez, Liad, et al.
Veröffentlicht: (2026)
Eluder dimension: localise it!
von: Bakhtiari, Alireza, et al.
Veröffentlicht: (2026)
von: Bakhtiari, Alireza, et al.
Veröffentlicht: (2026)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
The Real Price of Bandit Information in Multiclass Classification
von: Erez, Liad, et al.
Veröffentlicht: (2024)
von: Erez, Liad, et al.
Veröffentlicht: (2024)
Fast Rates for Bandit PAC Multiclass Classification
von: Erez, Liad, et al.
Veröffentlicht: (2024)
von: Erez, Liad, et al.
Veröffentlicht: (2024)
Rising Rested MAB with Linear Drift
von: Amichay, Omer, et al.
Veröffentlicht: (2025)
von: Amichay, Omer, et al.
Veröffentlicht: (2025)
Non-stochastic Bandits With Evolving Observations
von: Bar-On, Yogev, et al.
Veröffentlicht: (2024)
von: Bar-On, Yogev, et al.
Veröffentlicht: (2024)
Sample Complexity of Agnostic Multiclass Classification: Natarajan Dimension Strikes Back
von: Cohen, Alon, et al.
Veröffentlicht: (2025)
von: Cohen, Alon, et al.
Veröffentlicht: (2025)
Delay as Payoff in MAB
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2024)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2024)
Probably Approximately Precision and Recall Learning
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
Online Set Learning from Precision and Recall Feedback
von: Cohen, Lee, et al.
Veröffentlicht: (2026)
von: Cohen, Lee, et al.
Veröffentlicht: (2026)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
von: Erez, Liad, et al.
Veröffentlicht: (2022)
von: Erez, Liad, et al.
Veröffentlicht: (2022)
Truly No-Regret Learning in Constrained MDPs
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
How to Boost Any Loss Function
von: Nock, Richard, et al.
Veröffentlicht: (2024)
von: Nock, Richard, et al.
Veröffentlicht: (2024)
Learnability Gaps of Strategic Classification
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
von: Cohen, Lee, et al.
Veröffentlicht: (2024)
Swap Regret and Correlated Equilibria Beyond Normal-Form Games
von: Arunachaleswaran, Eshwar Ram, et al.
Veröffentlicht: (2025)
von: Arunachaleswaran, Eshwar Ram, et al.
Veröffentlicht: (2025)
Solving Robust MDPs through No-Regret Dynamics
von: Guha, Etash Kumar
Veröffentlicht: (2023)
von: Guha, Etash Kumar
Veröffentlicht: (2023)
No-Regret Reinforcement Learning in Smooth MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Online Weighted Paging with Unknown Weights
von: Levy, Orin, et al.
Veröffentlicht: (2024)
von: Levy, Orin, et al.
Veröffentlicht: (2024)
A Characterization of Semi-Supervised Adversarially-Robust PAC Learnability
von: Attias, Idan, et al.
Veröffentlicht: (2022)
von: Attias, Idan, et al.
Veröffentlicht: (2022)
Collaborating in Multi-Armed Bandits with Strategic Agents
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
von: Barnea, Idan, et al.
Veröffentlicht: (2026)
Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
von: Dann, Christoph, et al.
Veröffentlicht: (2026)
von: Dann, Christoph, et al.
Veröffentlicht: (2026)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
von: Li, Mingyi, et al.
Veröffentlicht: (2026)
von: Li, Mingyi, et al.
Veröffentlicht: (2026)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
Sample Complexity Characterization for Linear Contextual MDPs
von: Deng, Junze, et al.
Veröffentlicht: (2024)
von: Deng, Junze, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
von: Cassel, Asaf, et al.
Veröffentlicht: (2024) -
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026) -
Optimal Regret for Policy Optimization in Contextual Bandits
von: Levy, Orin, et al.
Veröffentlicht: (2026) -
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
von: Levy, Orin, et al.
Veröffentlicht: (2025) -
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
von: Barnea, Idan, et al.
Veröffentlicht: (2026)