Inference with the Upper Confidence Bound Algorithm
Fuente:
arXiv
Saved in:
| Main Authors: | Khamaru, Koulik, Zhang, Cun-Hui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024)
by: Khamaru, Koulik
Published: (2024)
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
Semi-parametric inference based on adaptively collected data
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Distribution-Free Confidence Ellipsoids for Ridge Regression with PAC Bounds
by: Szentpéteri, Szabolcs, et al.
Published: (2026)
by: Szentpéteri, Szabolcs, et al.
Published: (2026)
Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation
by: Sengupta, Saikat, et al.
Published: (2025)
by: Sengupta, Saikat, et al.
Published: (2025)
Efficient Inference after Directionally Stable Adaptive Experiments
by: Shen, Zikai, et al.
Published: (2026)
by: Shen, Zikai, et al.
Published: (2026)
From Data-Driven to Purpose-Driven Artificial Intelligence: Systems Thinking for Data-Analytic Automation of Patient Care
by: Anadria, Daniel, et al.
Published: (2025)
by: Anadria, Daniel, et al.
Published: (2025)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
by: Yan, Shunxing, et al.
Published: (2026)
by: Yan, Shunxing, et al.
Published: (2026)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
On Lai's Upper Confidence Bound in Multi-Armed Bandits
by: Ren, Huachen, et al.
Published: (2024)
by: Ren, Huachen, et al.
Published: (2024)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
by: Mohamed, Mimoun, et al.
Published: (2023)
by: Mohamed, Mimoun, et al.
Published: (2023)
An Upper Confidence Bound Approach to Estimating the Maximum Mean
by: Kun, Zhang, et al.
Published: (2024)
by: Kun, Zhang, et al.
Published: (2024)
Sampling Decisions
by: Chertkov, Michael, et al.
Published: (2025)
by: Chertkov, Michael, et al.
Published: (2025)
Temporal Memory for Resource-Constrained Agents: Continual Learning via Stochastic Compress-Add-Smooth
by: Chertkov, Michael
Published: (2026)
by: Chertkov, Michael
Published: (2026)
Adaptive Path Integral Diffusion: AdaPID
by: Chertkov, Michael, et al.
Published: (2025)
by: Chertkov, Michael, et al.
Published: (2025)
Generative Stochastic Optimal Transport: Guided Harmonic Path-Integral Diffusion
by: Chertkov, Michael
Published: (2025)
by: Chertkov, Michael
Published: (2025)
Analytic Bridge Diffusions for Controlled Path Generation
by: Chertkov, Michael
Published: (2026)
by: Chertkov, Michael
Published: (2026)
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026)
by: Nguyen, Minh
Published: (2026)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
by: Mustafi, Aratrika, et al.
Published: (2026)
by: Mustafi, Aratrika, et al.
Published: (2026)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
by: Kitaoka, Akira
Published: (2025)
by: Kitaoka, Akira
Published: (2025)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
by: He, Yixuan, et al.
Published: (2025)
by: He, Yixuan, et al.
Published: (2025)
FraPPE: Fast and Efficient Preference-based Pure Exploration
by: Das, Udvas, et al.
Published: (2025)
by: Das, Udvas, et al.
Published: (2025)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
by: Bareilles, Gilles, et al.
Published: (2023)
by: Bareilles, Gilles, et al.
Published: (2023)
Decentralized Upper Confidence Bound Algorithms for Homogeneous Multi-Agent Multi-Armed Bandits
by: Zhu, Jingxuan, et al.
Published: (2021)
by: Zhu, Jingxuan, et al.
Published: (2021)
Hierarchical Upper Confidence Bounds for Constrained Online Learning
by: Baheri, Ali
Published: (2024)
by: Baheri, Ali
Published: (2024)
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
Long-Context Linear System Identification
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
Finite-Sample Identification of Linear Regression Models with Residual-Permuted Sums
by: Szentpéteri, Szabolcs, et al.
Published: (2024)
by: Szentpéteri, Szabolcs, et al.
Published: (2024)
Rate-Optimal Non-Asymptotics for the Quadratic Prediction Error Method
by: Stamouli, Charis, et al.
Published: (2024)
by: Stamouli, Charis, et al.
Published: (2024)
Dual Unscented Kalman Filter Architecture for Sensor Fusion in Water Networks Leak Localization
by: Romero-Ben, Luis, et al.
Published: (2024)
by: Romero-Ben, Luis, et al.
Published: (2024)
Sample Complexity of the Sign-Perturbed Sums Identification Method: Scalar Case
by: Szentpéteri, Szabolcs, et al.
Published: (2024)
by: Szentpéteri, Szabolcs, et al.
Published: (2024)
Provably Bounding Neural Network Preimages
by: Kotha, Suhas, et al.
Published: (2023)
by: Kotha, Suhas, et al.
Published: (2023)
Similar Items
-
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024) -
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
by: Han, Qiyang, et al.
Published: (2024) -
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025) -
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025) -
Semi-parametric inference based on adaptively collected data
by: Lin, Licong, et al.
Published: (2023)