Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halder, Budhaditya, Sengupta, Ishan, Chowdhury, Koustav, Khamaru, Koulik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stable Thompson Sampling: Valid Inference via Variance Inflation
von: Halder, Budhaditya, et al.
Veröffentlicht: (2025)
von: Halder, Budhaditya, et al.
Veröffentlicht: (2025)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
von: Khamaru, Koulik
Veröffentlicht: (2024)
von: Khamaru, Koulik
Veröffentlicht: (2024)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
Bandit Simulation for Average Reward Inference
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation
von: Sengupta, Saikat, et al.
Veröffentlicht: (2025)
von: Sengupta, Saikat, et al.
Veröffentlicht: (2025)
Inference with the Upper Confidence Bound Algorithm
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
On the Effect of Regularization in Policy Mirror Descent
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
von: Sun, Haoyuan, et al.
Veröffentlicht: (2023)
von: Sun, Haoyuan, et al.
Veröffentlicht: (2023)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
Semi-parametric inference based on adaptively collected data
von: Lin, Licong, et al.
Veröffentlicht: (2023)
von: Lin, Licong, et al.
Veröffentlicht: (2023)
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
von: Han, Qiyang, et al.
Veröffentlicht: (2024)
von: Han, Qiyang, et al.
Veröffentlicht: (2024)
Uncertainty Quantification With Multiple Sources
von: Ying, Mufang, et al.
Veröffentlicht: (2024)
von: Ying, Mufang, et al.
Veröffentlicht: (2024)
Efficient Inference after Directionally Stable Adaptive Experiments
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
von: Shen, Zikai, et al.
Veröffentlicht: (2026)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
Consequences of Kernel Regularity for Bandit Optimization
von: Lee, Madison, et al.
Veröffentlicht: (2025)
von: Lee, Madison, et al.
Veröffentlicht: (2025)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
von: Karnik, Santhosh, et al.
Veröffentlicht: (2024)
von: Karnik, Santhosh, et al.
Veröffentlicht: (2024)
Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels
von: Yang, Jia-Qi, et al.
Veröffentlicht: (2025)
von: Yang, Jia-Qi, et al.
Veröffentlicht: (2025)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2026)
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2026)
On the Regularity and Fairness of Combinatorial Multi-Armed Bandit
von: Wu, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyi, et al.
Veröffentlicht: (2025)
Enhancing Robust Fairness via Confusional Spectral Regularization
von: Jin, Gaojie, et al.
Veröffentlicht: (2025)
von: Jin, Gaojie, et al.
Veröffentlicht: (2025)
Mirror Descent Actor Critic via Bounded Advantage Learning
von: Iwaki, Ryo
Veröffentlicht: (2025)
von: Iwaki, Ryo
Veröffentlicht: (2025)
Robust Implicit Regularization via Weight Normalization
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
Semi-Discrete Optimal Transport: Nearly Minimax Estimation With Stochastic Gradient Descent and Adaptive Entropic Regularization
von: Genans, Ferdinand, et al.
Veröffentlicht: (2024)
von: Genans, Ferdinand, et al.
Veröffentlicht: (2024)
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits
von: Huang, Ziyi, et al.
Veröffentlicht: (2024)
von: Huang, Ziyi, et al.
Veröffentlicht: (2024)
Variational Online Mirror Descent for Robust Learning in Schrödinger Bridge
von: Han, Dong-Sig, et al.
Veröffentlicht: (2025)
von: Han, Dong-Sig, et al.
Veröffentlicht: (2025)
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
von: Li, Tong, et al.
Veröffentlicht: (2026)
von: Li, Tong, et al.
Veröffentlicht: (2026)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
von: Zhao, Heyang, et al.
Veröffentlicht: (2024)
von: Zhao, Heyang, et al.
Veröffentlicht: (2024)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2026)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2026)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
von: Ye, Qiang
Veröffentlicht: (2024)
von: Ye, Qiang
Veröffentlicht: (2024)
Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Improving Robustness In Sparse Autoencoders via Masked Regularization
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
On Generalization and Regularization via Wasserstein Distributionally Robust Optimization
von: Wu, Qinyu, et al.
Veröffentlicht: (2022)
von: Wu, Qinyu, et al.
Veröffentlicht: (2022)
Passive Inference Attacks on Split Learning via Adversarial Regularization
von: Zhu, Xiaochen, et al.
Veröffentlicht: (2023)
von: Zhu, Xiaochen, et al.
Veröffentlicht: (2023)
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization
von: Kumar, Ramnath, et al.
Veröffentlicht: (2023)
von: Kumar, Ramnath, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Stable Thompson Sampling: Valid Inference via Variance Inflation
von: Halder, Budhaditya, et al.
Veröffentlicht: (2025) -
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
von: Praharaj, Samya, et al.
Veröffentlicht: (2025) -
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
von: Khamaru, Koulik
Veröffentlicht: (2024) -
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025) -
Bandit Simulation for Average Reward Inference
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)