Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Haolin, Mhammedi, Zakaria, Wei, Chen-Yu, Zimmert, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
by: Liu, Haolin, et al.
Published: (2025)
by: Liu, Haolin, et al.
Published: (2025)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023)
by: Mhammedi, Zakaria, et al.
Published: (2023)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Decision Making in Hybrid Environments: A Model Aggregation Approach
by: Liu, Haolin, et al.
Published: (2025)
by: Liu, Haolin, et al.
Published: (2025)
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
by: Masoudian, Saeed, et al.
Published: (2023)
by: Masoudian, Saeed, et al.
Published: (2023)
Online Convex Optimization with a Separation Oracle
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Incentive-compatible Bandits: Importance Weighting No More
by: Zimmert, Julian, et al.
Published: (2024)
by: Zimmert, Julian, et al.
Published: (2024)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024)
by: Cutkosky, Ashok, et al.
Published: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Non-stationary Bandit Convex Optimization: A Comprehensive Study
by: Liu, Xiaoqi, et al.
Published: (2025)
by: Liu, Xiaoqi, et al.
Published: (2025)
A Model Selection Approach for Corruption Robust Reinforcement Learning
by: Wei, Chen-Yu, et al.
Published: (2021)
by: Wei, Chen-Yu, et al.
Published: (2021)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Optimal cross-learning for contextual bandits with unknown context distributions
by: Schneider, Jon, et al.
Published: (2024)
by: Schneider, Jon, et al.
Published: (2024)
The Power of Resets in Online Reinforcement Learning
by: Mhammedi, Zakaria, et al.
Published: (2024)
by: Mhammedi, Zakaria, et al.
Published: (2024)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Adaptive Matrix Online Learning through Smoothing with Guarantees for Nonsmooth Nonconvex Optimization
by: Jiang, Ruichen, et al.
Published: (2026)
by: Jiang, Ruichen, et al.
Published: (2026)
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025)
by: Foster, Dylan J., et al.
Published: (2025)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)
by: Kim, Jungtaek, et al.
Published: (2026)
A Bayesian Interpretation of Adaptive Low-Rank Adaptation
by: Chen, Haolin, et al.
Published: (2024)
by: Chen, Haolin, et al.
Published: (2024)
Efficient Generalized Low-Rank Tensor Contextual Bandits
by: Yi, Qianxin, et al.
Published: (2023)
by: Yi, Qianxin, et al.
Published: (2023)
Adversarial Bandits with Multi-User Delayed Feedback: Theory and Application
by: Li, Yandi, et al.
Published: (2023)
by: Li, Yandi, et al.
Published: (2023)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
by: Huang, Ruiquan, et al.
Published: (2026)
by: Huang, Ruiquan, et al.
Published: (2026)
Addressing Finite-Horizon MDPs via Low-Rank Tensor Value Approximation
by: Rozada, Sergio, et al.
Published: (2025)
by: Rozada, Sergio, et al.
Published: (2025)
Performative Prediction with Bandit Feedback: Learning through Reparameterization
by: Chen, Yatong, et al.
Published: (2023)
by: Chen, Yatong, et al.
Published: (2023)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)
by: Gangrade, Aditya, et al.
Published: (2022)
Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback
by: Lee, Seong Jin, et al.
Published: (2024)
by: Lee, Seong Jin, et al.
Published: (2024)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
by: Lee, Minjae, et al.
Published: (2025)
by: Lee, Minjae, et al.
Published: (2025)
Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions
by: Oh, Youngmin
Published: (2026)
by: Oh, Youngmin
Published: (2026)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
by: Kang, Yue, et al.
Published: (2025)
by: Kang, Yue, et al.
Published: (2025)
Similar Items
-
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
by: Liu, Haolin, et al.
Published: (2025) -
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024) -
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023) -
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026) -
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024)