Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jing, Zhao, Peng, Zhou, Zhi-Hua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Efficient Methods for Non-stationary Online Learning
by: Zhao, Peng, et al.
Published: (2023)
by: Zhao, Peng, et al.
Published: (2023)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Non-Parametric Rehearsal Learning via Conditional Mean Embeddings
by: Du, Wen-Bo, et al.
Published: (2026)
by: Du, Wen-Bo, et al.
Published: (2026)
Non-stationary Bandit Convex Optimization: A Comprehensive Study
by: Liu, Xiaoqi, et al.
Published: (2025)
by: Liu, Xiaoqi, et al.
Published: (2025)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
by: Shrestha, Aayam, et al.
Published: (2020)
by: Shrestha, Aayam, et al.
Published: (2020)
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Non-stationary Delayed Online Convex Optimization: From Full-information to Bandit Setting
by: Wan, Yuanyu, et al.
Published: (2023)
by: Wan, Yuanyu, et al.
Published: (2023)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
On Pareto Optimality for Parametric Choice Bandits
by: Zuo, Jierui, et al.
Published: (2025)
by: Zuo, Jierui, et al.
Published: (2025)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Extended UCB Policies for Multi-armed Bandit Problems
by: Liu, Keqin, et al.
Published: (2011)
by: Liu, Keqin, et al.
Published: (2011)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Parametric Prior Mapping Framework for Non-stationary Probabilistic Time Series Forecasting
by: Li, Jinglin, et al.
Published: (2026)
by: Li, Jinglin, et al.
Published: (2026)
Adaptivity and Non-stationarity: Problem-dependent Dynamic Regret for Online Convex Optimization
by: Zhao, Peng, et al.
Published: (2021)
by: Zhao, Peng, et al.
Published: (2021)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
by: Kang, Yue, et al.
Published: (2025)
by: Kang, Yue, et al.
Published: (2025)
Partition Tree Weighting for Non-Stationary Stochastic Bandits
by: Veness, Joel, et al.
Published: (2025)
by: Veness, Joel, et al.
Published: (2025)
Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits
by: Werge, Nicklas, et al.
Published: (2023)
by: Werge, Nicklas, et al.
Published: (2023)
Non-Stationary Dueling Bandits Under a Weighted Borda Criterion
by: Suk, Joe, et al.
Published: (2024)
by: Suk, Joe, et al.
Published: (2024)
Revisiting Matrix Sketching in Linear Bandits: Achieving Sublinear Regret via Dyadic Block Sketching
by: Wen, Dongxie, et al.
Published: (2024)
by: Wen, Dongxie, et al.
Published: (2024)
Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability
by: Zhang, Yu-Jie, et al.
Published: (2025)
by: Zhang, Yu-Jie, et al.
Published: (2025)
Diversity-Preserving K-Armed Bandits, Revisited
by: Hadiji, Hédi, et al.
Published: (2020)
by: Hadiji, Hédi, et al.
Published: (2020)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
A Simple, Optimal and Efficient Algorithm for Online Exp-Concave Optimization
by: Wang, Yi-Han, et al.
Published: (2025)
by: Wang, Yi-Han, et al.
Published: (2025)
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Mixture of Online and Offline Experts for Non-stationary Time Series
by: Zhao, Zhilin, et al.
Published: (2022)
by: Zhao, Zhilin, et al.
Published: (2022)
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
by: Wang, Xiaoshuang, et al.
Published: (2025)
by: Wang, Xiaoshuang, et al.
Published: (2025)
Robust Length Prediction: A Perspective from Heavy-Tailed Prompt-Conditioned Distributions
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Linear Contextual Bandits with Hybrid Payoff: Revisited
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
Similar Items
-
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024) -
Efficient Methods for Non-stationary Online Learning
by: Zhao, Peng, et al.
Published: (2023) -
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
by: Li, Long-Fei, et al.
Published: (2024) -
Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update
by: Wang, Jing, et al.
Published: (2025) -
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)