Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kash, Ian A., Reyzin, Lev, Yu, Zishun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning-Augmented Algorithms for Boolean Satisfiability
di: Attias, Idan, et al.
Pubblicazione: (2025)
di: Attias, Idan, et al.
Pubblicazione: (2025)
Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
di: Thoppe, Gugan, et al.
Pubblicazione: (2026)
di: Thoppe, Gugan, et al.
Pubblicazione: (2026)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
di: Lyu, Lixing, et al.
Pubblicazione: (2025)
di: Lyu, Lixing, et al.
Pubblicazione: (2025)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
Contradiction Graphs Determine VC Dimension
di: Campbell, Jesse, et al.
Pubblicazione: (2026)
di: Campbell, Jesse, et al.
Pubblicazione: (2026)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
di: Liu, Haolin, et al.
Pubblicazione: (2024)
di: Liu, Haolin, et al.
Pubblicazione: (2024)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
di: Grand-Clément, Julien, et al.
Pubblicazione: (2023)
di: Grand-Clément, Julien, et al.
Pubblicazione: (2023)
Non-adaptive Learning of Random Hypergraphs with Queries
di: Austhof, Bethany, et al.
Pubblicazione: (2025)
di: Austhof, Bethany, et al.
Pubblicazione: (2025)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
di: Ito, Shinji, et al.
Pubblicazione: (2025)
di: Ito, Shinji, et al.
Pubblicazione: (2025)
On the Hardness of Learning Regular Expressions
di: Attias, Idan, et al.
Pubblicazione: (2025)
di: Attias, Idan, et al.
Pubblicazione: (2025)
Choice of Scoring Rules for Indirect Elicitation of Properties with Parametric Assumptions
di: Hu, Lingfang, et al.
Pubblicazione: (2025)
di: Hu, Lingfang, et al.
Pubblicazione: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
di: Viano, Luca, et al.
Pubblicazione: (2024)
di: Viano, Luca, et al.
Pubblicazione: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
di: Hong, Kihyuk, et al.
Pubblicazione: (2024)
di: Hong, Kihyuk, et al.
Pubblicazione: (2024)
Data-Driven Adversarial Online Control for Unknown Linear Systems
di: Liu, Zishun, et al.
Pubblicazione: (2023)
di: Liu, Zishun, et al.
Pubblicazione: (2023)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
di: van Erven, Tim, et al.
Pubblicazione: (2025)
di: van Erven, Tim, et al.
Pubblicazione: (2025)
Nearly-Optimal Algorithm for Adversarial Kernelized Bandits
di: Iwazaki, Shogo
Pubblicazione: (2026)
di: Iwazaki, Shogo
Pubblicazione: (2026)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
di: Cassel, Asaf, et al.
Pubblicazione: (2024)
di: Cassel, Asaf, et al.
Pubblicazione: (2024)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
di: Lu, Michael, et al.
Pubblicazione: (2024)
di: Lu, Michael, et al.
Pubblicazione: (2024)
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
di: Wang, Jing, et al.
Pubblicazione: (2026)
di: Wang, Jing, et al.
Pubblicazione: (2026)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
di: Manivannan, Sanjeev, et al.
Pubblicazione: (2026)
di: Manivannan, Sanjeev, et al.
Pubblicazione: (2026)
A Jointly Efficient and Optimal Algorithm for Heteroskedastic Generalized Linear Bandits with Adversarial Corruptions
di: Kim, Sanghwa, et al.
Pubblicazione: (2026)
di: Kim, Sanghwa, et al.
Pubblicazione: (2026)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
di: Eshwar, S. R.
Pubblicazione: (2025)
di: Eshwar, S. R.
Pubblicazione: (2025)
Learning Adversarial MDPs with Stochastic Hard Constraints
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
DISCO: An End-to-End Bandit Framework for Personalised Discount Allocation
di: Zhang, Jason Shuo, et al.
Pubblicazione: (2024)
di: Zhang, Jason Shuo, et al.
Pubblicazione: (2024)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
di: Goyal, Tanmay, et al.
Pubblicazione: (2025)
di: Goyal, Tanmay, et al.
Pubblicazione: (2025)
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
di: Liu, Haolin, et al.
Pubblicazione: (2025)
di: Liu, Haolin, et al.
Pubblicazione: (2025)
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
di: Satheesh, Anirudh, et al.
Pubblicazione: (2026)
di: Satheesh, Anirudh, et al.
Pubblicazione: (2026)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
di: Shah, Anvay, et al.
Pubblicazione: (2026)
di: Shah, Anvay, et al.
Pubblicazione: (2026)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
di: Tsuchiya, Taira, et al.
Pubblicazione: (2025)
di: Tsuchiya, Taira, et al.
Pubblicazione: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2026)
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2026)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
di: Mortensen, Oliver, et al.
Pubblicazione: (2025)
di: Mortensen, Oliver, et al.
Pubblicazione: (2025)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
di: Di, Qiwei, et al.
Pubblicazione: (2024)
di: Di, Qiwei, et al.
Pubblicazione: (2024)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
di: Banerjee, Debangshu, et al.
Pubblicazione: (2023)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2023)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
di: Lancewicki, Tal, et al.
Pubblicazione: (2025)
di: Lancewicki, Tal, et al.
Pubblicazione: (2025)
Efficient and Adaptive Posterior Sampling Algorithms for Bandits
di: Hu, Bingshan, et al.
Pubblicazione: (2024)
di: Hu, Bingshan, et al.
Pubblicazione: (2024)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
Representative Action Selection for Large Action Space: From Bandits to MDPs
di: Zhou, Quan, et al.
Pubblicazione: (2025)
di: Zhou, Quan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Learning-Augmented Algorithms for Boolean Satisfiability
di: Attias, Idan, et al.
Pubblicazione: (2025) -
Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
di: Thoppe, Gugan, et al.
Pubblicazione: (2026) -
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
di: Lyu, Lixing, et al.
Pubblicazione: (2025) -
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
di: Li, Long-Fei, et al.
Pubblicazione: (2024) -
Contradiction Graphs Determine VC Dimension
di: Campbell, Jesse, et al.
Pubblicazione: (2026)