Non-stationary Bandit Convex Optimization: A Comprehensive Study
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaoqi, Baudry, Dorian, Zimmert, Julian, Rebeschini, Patrick, Akhavan, Arya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Concordant Perturbations for Linear Bandits
by: Lévy, Lucas, et al.
Published: (2025)
by: Lévy, Lucas, et al.
Published: (2025)
Best-of-Both Worlds for linear contextual bandits with paid observations
by: Boyer, Nathan, et al.
Published: (2025)
by: Boyer, Nathan, et al.
Published: (2025)
A General Recipe for the Analysis of Randomized Multi-Armed Bandit Algorithms
by: Baudry, Dorian, et al.
Published: (2023)
by: Baudry, Dorian, et al.
Published: (2023)
A Perturbation Approach to Unconstrained Linear Bandits
by: Jacobsen, Andrew, et al.
Published: (2026)
by: Jacobsen, Andrew, et al.
Published: (2026)
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
by: Masoudian, Saeed, et al.
Published: (2023)
by: Masoudian, Saeed, et al.
Published: (2023)
Incentive-compatible Bandits: Importance Weighting No More
by: Zimmert, Julian, et al.
Published: (2024)
by: Zimmert, Julian, et al.
Published: (2024)
Non-stationary Delayed Online Convex Optimization: From Full-information to Bandit Setting
by: Wan, Yuanyu, et al.
Published: (2023)
by: Wan, Yuanyu, et al.
Published: (2023)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Optimal cross-learning for contextual bandits with unknown context distributions
by: Schneider, Jon, et al.
Published: (2024)
by: Schneider, Jon, et al.
Published: (2024)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Robust Gradient Descent for Phase Retrieval
by: Buna, Alex, et al.
Published: (2024)
by: Buna, Alex, et al.
Published: (2024)
Decision Making in Hybrid Environments: A Model Aggregation Approach
by: Liu, Haolin, et al.
Published: (2025)
by: Liu, Haolin, et al.
Published: (2025)
Gradient-free stochastic optimization for additive models
by: Akhavan, Arya, et al.
Published: (2025)
by: Akhavan, Arya, et al.
Published: (2025)
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
by: Liu, Haolin, et al.
Published: (2025)
by: Liu, Haolin, et al.
Published: (2025)
A Model Selection Approach for Corruption Robust Reinforcement Learning
by: Wei, Chen-Yu, et al.
Published: (2021)
by: Wei, Chen-Yu, et al.
Published: (2021)
Sharp analysis of linear ensemble sampling
by: Akhavan, Arya, et al.
Published: (2026)
by: Akhavan, Arya, et al.
Published: (2026)
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Improved Regret for Bandit Convex Optimization with Delayed Feedback
by: Wan, Yuanyu, et al.
Published: (2024)
by: Wan, Yuanyu, et al.
Published: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
by: Wegel, Tobias, et al.
Published: (2025)
by: Wegel, Tobias, et al.
Published: (2025)
Bandit Convex Optimization with Gradient Prediction Adaptivity
by: Wang, Shuche, et al.
Published: (2026)
by: Wang, Shuche, et al.
Published: (2026)
A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
by: Alfano, Carlo, et al.
Published: (2023)
by: Alfano, Carlo, et al.
Published: (2023)
High-probability zeroth-order online convex optimisation beyond Euclidean geometry
by: Janz, David, et al.
Published: (2025)
by: Janz, David, et al.
Published: (2025)
Improved Dimension Dependence for Bandit Convex Optimization with Gradient Variations
by: Yu, Hang, et al.
Published: (2026)
by: Yu, Hang, et al.
Published: (2026)
Differentiable Cost-Parameterized Monge Map Estimators
by: Howard, Samuel, et al.
Published: (2024)
by: Howard, Samuel, et al.
Published: (2024)
Sample-Efficiency in Multi-Batch Reinforcement Learning: The Need for Dimension-Dependent Adaptivity
by: Johnson, Emmeran, et al.
Published: (2023)
by: Johnson, Emmeran, et al.
Published: (2023)
Bandit Convex Optimisation
by: Lattimore, Tor
Published: (2024)
by: Lattimore, Tor
Published: (2024)
Quantum Non-Linear Bandit Optimization
by: Siam, Zakaria Shams, et al.
Published: (2025)
by: Siam, Zakaria Shams, et al.
Published: (2025)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
A conversion theorem and minimax optimality for continuum contextual bandits
by: Akhavan, Arya, et al.
Published: (2024)
by: Akhavan, Arya, et al.
Published: (2024)
Contextual Dynamic Pricing with Heterogeneous Buyers
by: Lykouris, Thodoris, et al.
Published: (2025)
by: Lykouris, Thodoris, et al.
Published: (2025)
Stochastic Shortest Path with Sparse Adversarial Costs
by: Johnson, Emmeran, et al.
Published: (2025)
by: Johnson, Emmeran, et al.
Published: (2025)
Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis
by: Farghly, Tyler, et al.
Published: (2025)
by: Farghly, Tyler, et al.
Published: (2025)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
by: Vary, Simon, et al.
Published: (2024)
by: Vary, Simon, et al.
Published: (2024)
Improved learning rates in multi-unit uniform price auctions
by: Potfer, Marius, et al.
Published: (2025)
by: Potfer, Marius, et al.
Published: (2025)
Optimal High-Probability Regret for Online Convex Optimization with Two-Point Bandit Feedback
by: Ye, Haishan
Published: (2026)
by: Ye, Haishan
Published: (2026)
On the necessity of adaptive regularisation:Optimal anytime online learning on $\boldsymbol{\ell_p}$-balls
by: Johnson, Emmeran, et al.
Published: (2025)
by: Johnson, Emmeran, et al.
Published: (2025)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
by: Akgül, Abdullah, et al.
Published: (2025)
by: Akgül, Abdullah, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Similar Items
-
Self-Concordant Perturbations for Linear Bandits
by: Lévy, Lucas, et al.
Published: (2025) -
Best-of-Both Worlds for linear contextual bandits with paid observations
by: Boyer, Nathan, et al.
Published: (2025) -
A General Recipe for the Analysis of Randomized Multi-Armed Bandit Algorithms
by: Baudry, Dorian, et al.
Published: (2023) -
A Perturbation Approach to Unconstrained Linear Bandits
by: Jacobsen, Andrew, et al.
Published: (2026) -
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
by: Masoudian, Saeed, et al.
Published: (2023)