Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Long-Fei, Zhao, Peng, Zhou, Zhi-Hua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
by: Kang, Yue, et al.
Published: (2025)
by: Kang, Yue, et al.
Published: (2025)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
by: van Erven, Tim, et al.
Published: (2025)
by: van Erven, Tim, et al.
Published: (2025)
Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures
by: Flynn, Hamish, et al.
Published: (2023)
by: Flynn, Hamish, et al.
Published: (2023)
Exploratory Machine Learning with Unknown Unknowns
by: Zhao, Peng, et al.
Published: (2020)
by: Zhao, Peng, et al.
Published: (2020)
Improved Algorithms for Nash Welfare in Linear Bandits
by: Sarkar, Dhruv, et al.
Published: (2026)
by: Sarkar, Dhruv, et al.
Published: (2026)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions
by: Oh, Youngmin
Published: (2026)
by: Oh, Youngmin
Published: (2026)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)
by: Gangrade, Aditya, et al.
Published: (2022)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Linear Causal Bandits: Unknown Graph and Soft Interventions
by: Yan, Zirui, et al.
Published: (2024)
by: Yan, Zirui, et al.
Published: (2024)
Graph Feedback Bandits with Similar Arms
by: Qi, Han, et al.
Published: (2024)
by: Qi, Han, et al.
Published: (2024)
Adversarial Bandits with Multi-User Delayed Feedback: Theory and Application
by: Li, Yandi, et al.
Published: (2023)
by: Li, Yandi, et al.
Published: (2023)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)
by: Kim, Jungtaek, et al.
Published: (2026)
Differentially Private Linear Bandits with Partial Distributed Feedback
by: Li, Fengjiao, et al.
Published: (2022)
by: Li, Fengjiao, et al.
Published: (2022)
Heavy-tailed Linear Bandits: Adversarial Robustness, Best-of-both-worlds, and Beyond
by: Zhao, Canzhe, et al.
Published: (2025)
by: Zhao, Canzhe, et al.
Published: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
by: Goyal, Tanmay, et al.
Published: (2025)
by: Goyal, Tanmay, et al.
Published: (2025)
A Jointly Efficient and Optimal Algorithm for Heteroskedastic Generalized Linear Bandits with Adversarial Corruptions
by: Kim, Sanghwa, et al.
Published: (2026)
by: Kim, Sanghwa, et al.
Published: (2026)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Sparsity-Agnostic Linear Bandits with Adaptive Adversaries
by: Jin, Tianyuan, et al.
Published: (2024)
by: Jin, Tianyuan, et al.
Published: (2024)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
by: Liu, Haolin, et al.
Published: (2025)
by: Liu, Haolin, et al.
Published: (2025)
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits
by: Huang, Ziyi, et al.
Published: (2024)
by: Huang, Ziyi, et al.
Published: (2024)
Causal Bandit Over Unknown Graphs: Upper Confidence Bounds With Backdoor Adjustment
by: Zhao, Yijia, et al.
Published: (2025)
by: Zhao, Yijia, et al.
Published: (2025)
Linear Submodular Maximization with Bandit Feedback
by: Chen, Wenjing, et al.
Published: (2024)
by: Chen, Wenjing, et al.
Published: (2024)
Nearly-Optimal Algorithm for Adversarial Kernelized Bandits
by: Iwazaki, Shogo
Published: (2026)
by: Iwazaki, Shogo
Published: (2026)
Learning Kernel-Based MDPs from Episodic Preferential Feedback
by: Pavlovic, Nikola, et al.
Published: (2026)
by: Pavlovic, Nikola, et al.
Published: (2026)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Similar Items
-
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
by: Li, Long-Fei, et al.
Published: (2024) -
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024) -
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
by: Wang, Jing, et al.
Published: (2026) -
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024) -
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)