Saved in:
| Main Authors: | Praharaj, Samya, Chang, Chih-Yu, Khamaru, Koulik, Zhang, Kelly W. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2606.00913 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024)
by: Khamaru, Koulik
Published: (2024)
Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent
by: Halder, Budhaditya, et al.
Published: (2026)
by: Halder, Budhaditya, et al.
Published: (2026)
Inference with the Upper Confidence Bound Algorithm
by: Khamaru, Koulik, et al.
Published: (2024)
by: Khamaru, Koulik, et al.
Published: (2024)
Stable Thompson Sampling: Valid Inference via Variance Inflation
by: Halder, Budhaditya, et al.
Published: (2025)
by: Halder, Budhaditya, et al.
Published: (2025)
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Semi-parametric inference based on adaptively collected data
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Uncertainty Quantification With Multiple Sources
by: Ying, Mufang, et al.
Published: (2024)
by: Ying, Mufang, et al.
Published: (2024)
Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation
by: Sengupta, Saikat, et al.
Published: (2025)
by: Sengupta, Saikat, et al.
Published: (2025)
Efficient Inference after Directionally Stable Adaptive Experiments
by: Shen, Zikai, et al.
Published: (2026)
by: Shen, Zikai, et al.
Published: (2026)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
PICS: A sequential approach to obtain optimal designs for non-linear models leveraging closed-form solutions for faster convergence
by: Ghosh, Suvrojit, et al.
Published: (2024)
by: Ghosh, Suvrojit, et al.
Published: (2024)
Transductive Reward Inference on Graph
by: Qu, Bohao, et al.
Published: (2024)
by: Qu, Bohao, et al.
Published: (2024)
BanditQ: Fair Bandits with Guaranteed Rewards
by: Sinha, Abhishek
Published: (2023)
by: Sinha, Abhishek
Published: (2023)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2023)
by: Hong, Yige, et al.
Published: (2023)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
by: Kang, Yue, et al.
Published: (2025)
by: Kang, Yue, et al.
Published: (2025)
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits
by: Huang, Ziyi, et al.
Published: (2024)
by: Huang, Ziyi, et al.
Published: (2024)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
by: Ye, Chenlu, et al.
Published: (2025)
by: Ye, Chenlu, et al.
Published: (2025)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
by: Chang, Xiangyu, et al.
Published: (2022)
by: Chang, Xiangyu, et al.
Published: (2022)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
by: Qin, Hao, et al.
Published: (2023)
by: Qin, Hao, et al.
Published: (2023)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
by: Roch, Zachary, et al.
Published: (2025)
by: Roch, Zachary, et al.
Published: (2025)
Low-rank Matrix Bandits with Heavy-tailed Rewards
by: Kang, Yue, et al.
Published: (2024)
by: Kang, Yue, et al.
Published: (2024)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Transfer in Sequential Multi-armed Bandits via Reward Samples
by: R, Rahul N, et al.
Published: (2024)
by: R, Rahul N, et al.
Published: (2024)
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Supervised Reward Inference
by: Schwarzer, Will, et al.
Published: (2025)
by: Schwarzer, Will, et al.
Published: (2025)
Impatient Bandits: Optimizing for the Long-Term Without Delay
by: Zhang, Kelly W., et al.
Published: (2025)
by: Zhang, Kelly W., et al.
Published: (2025)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
A Modularized Framework for Piecewise-Stationary Restless Bandits
by: Li, Kuan-Ta, et al.
Published: (2026)
by: Li, Kuan-Ta, et al.
Published: (2026)
Diminishing Exploration: A Minimalist Approach to Piecewise Stationary Multi-Armed Bandits
by: Li, Kuan-Ta, et al.
Published: (2024)
by: Li, Kuan-Ta, et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
by: Wan, Yi, et al.
Published: (2024)
by: Wan, Yi, et al.
Published: (2024)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Similar Items
-
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025) -
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025) -
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024) -
Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent
by: Halder, Budhaditya, et al.
Published: (2026) -
Inference with the Upper Confidence Bound Algorithm
by: Khamaru, Koulik, et al.
Published: (2024)