Optimism Stabilizes Thompson Sampling for Adaptive Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Shunxing, Zhong, Han |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
by: Roy, Somjit, et al.
Published: (2026)
by: Roy, Somjit, et al.
Published: (2026)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026)
by: Nguyen, Minh
Published: (2026)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026)
by: Mustafi, Aratrika, et al.
Published: (2026)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
by: Mohamed, Mimoun, et al.
Published: (2023)
by: Mohamed, Mimoun, et al.
Published: (2023)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
by: Kitaoka, Akira
Published: (2025)
by: Kitaoka, Akira
Published: (2025)
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
by: He, Yixuan, et al.
Published: (2025)
by: He, Yixuan, et al.
Published: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
FraPPE: Fast and Efficient Preference-based Pure Exploration
by: Das, Udvas, et al.
Published: (2025)
by: Das, Udvas, et al.
Published: (2025)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
by: Bareilles, Gilles, et al.
Published: (2023)
by: Bareilles, Gilles, et al.
Published: (2023)
MINTS: Minimalist Thompson Sampling
by: Wang, Kaizheng
Published: (2026)
by: Wang, Kaizheng
Published: (2026)
Robust Assortment Optimization from Observational Data
by: Lu, Miao, et al.
Published: (2026)
by: Lu, Miao, et al.
Published: (2026)
Learning an Optimal Assortment Policy under Observational Data
by: Han, Yuxuan, et al.
Published: (2025)
by: Han, Yuxuan, et al.
Published: (2025)
Stochastic Optimization with Optimal Importance Sampling
by: Aolaritei, Liviu, et al.
Published: (2025)
by: Aolaritei, Liviu, et al.
Published: (2025)
Online Inference of Constrained Optimization: Primal-Dual Optimality and Sequential Quadratic Programming
by: Gao, Yihang, et al.
Published: (2025)
by: Gao, Yihang, et al.
Published: (2025)
An Improved Analysis of Langevin Algorithms with Prior Diffusion for Non-Log-Concave Sampling
by: Huang, Xunpeng, et al.
Published: (2024)
by: Huang, Xunpeng, et al.
Published: (2024)
On the Sample Complexity of Set Membership Estimation for Linear Systems with Disturbances Bounded by Convex Sets
by: Xu, Haonan, et al.
Published: (2024)
by: Xu, Haonan, et al.
Published: (2024)
Probabilistic Guarantees of Stochastic Recursive Gradient in Non-Convex Finite Sum Problems
by: Zhong, Yanjie, et al.
Published: (2024)
by: Zhong, Yanjie, et al.
Published: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
Early Stopping in Contextual Bandits and Inferences
by: Cui, Zihan
Published: (2025)
by: Cui, Zihan
Published: (2025)
Analysis of Thompson Sampling for Controlling Unknown Linear Diffusion Processes
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
Ensemble-Conditional Gaussian Processes (Ens-CGP): Representation, Geometry, and Inference
by: Ravela, Sai, et al.
Published: (2026)
by: Ravela, Sai, et al.
Published: (2026)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
by: Li, Gen, et al.
Published: (2021)
by: Li, Gen, et al.
Published: (2021)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
by: Agrawalla, Bhavya, et al.
Published: (2023)
by: Agrawalla, Bhavya, et al.
Published: (2023)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
by: Li, Gen, et al.
Published: (2020)
by: Li, Gen, et al.
Published: (2020)
Long-time dynamics and universality of nonconvex gradient descent
by: Han, Qiyang
Published: (2025)
by: Han, Qiyang
Published: (2025)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
Conformal Prediction in The Loop: A Feedback-Based Uncertainty Model for Trajectory Optimization
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Analytic Bridge Diffusions for Controlled Path Generation
by: Chertkov, Michael
Published: (2026)
by: Chertkov, Michael
Published: (2026)
Conformal Prediction-Driven Adaptive Sampling for Digital Water Twins
by: Homaei, Mohammadhossein, et al.
Published: (2025)
by: Homaei, Mohammadhossein, et al.
Published: (2025)
A Graphical Global Optimization Framework for Parameter Estimation of Statistical Models with Nonconvex Regularization Functions
by: Davarnia, Danial, et al.
Published: (2025)
by: Davarnia, Danial, et al.
Published: (2025)
Inference with the Upper Confidence Bound Algorithm
by: Khamaru, Koulik, et al.
Published: (2024)
by: Khamaru, Koulik, et al.
Published: (2024)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
Inference-Time Alignment for Diffusion Models via Variationally Stable Doob's Matching
by: Chang, Jinyuan, et al.
Published: (2026)
by: Chang, Jinyuan, et al.
Published: (2026)
Similar Items
-
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025) -
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
by: Roy, Somjit, et al.
Published: (2026) -
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026) -
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026) -
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026)