Upper Counterfactual Confidence Bounds: a New Optimism Principle for Contextual Bandits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yunbei, Zeevi, Assaf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bayesian Design Principles for Frequentist Sequential Learning
von: Xu, Yunbei, et al.
Veröffentlicht: (2023)
von: Xu, Yunbei, et al.
Veröffentlicht: (2023)
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
von: Chen, Fan, et al.
Veröffentlicht: (2024)
von: Chen, Fan, et al.
Veröffentlicht: (2024)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
An Upper Confidence Bound Approach to Estimating the Maximum Mean
von: Kun, Zhang, et al.
Veröffentlicht: (2024)
von: Kun, Zhang, et al.
Veröffentlicht: (2024)
On Lai's Upper Confidence Bound in Multi-Armed Bandits
von: Ren, Huachen, et al.
Veröffentlicht: (2024)
von: Ren, Huachen, et al.
Veröffentlicht: (2024)
Batched Nonparametric Contextual Bandits
von: Jiang, Rong, et al.
Veröffentlicht: (2024)
von: Jiang, Rong, et al.
Veröffentlicht: (2024)
Inference with the Upper Confidence Bound Algorithm
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
Pointwise Generalization in Deep Neural Networks
von: Li, Shaojie, et al.
Veröffentlicht: (2026)
von: Li, Shaojie, et al.
Veröffentlicht: (2026)
Gaussian Process Upper Confidence Bounds in Distributed Point Target Tracking over Wireless Sensor Networks
von: Liu, Xingchi, et al.
Veröffentlicht: (2024)
von: Liu, Xingchi, et al.
Veröffentlicht: (2024)
Transfer Learning for Contextual Multi-armed Bandits
von: Cai, Changxiao, et al.
Veröffentlicht: (2022)
von: Cai, Changxiao, et al.
Veröffentlicht: (2022)
Navigating Sparsities in High-Dimensional Linear Contextual Bandits
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
von: Réveillard, William, et al.
Veröffentlicht: (2025)
von: Réveillard, William, et al.
Veröffentlicht: (2025)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
Revisiting Optimism and Model Complexity in the Wake of Overparameterized Machine Learning
von: Patil, Pratik, et al.
Veröffentlicht: (2024)
von: Patil, Pratik, et al.
Veröffentlicht: (2024)
Asymptotic Optimism for Tensor Regression Models with Applications to Neural Network Compression
von: Shi, Haoming, et al.
Veröffentlicht: (2026)
von: Shi, Haoming, et al.
Veröffentlicht: (2026)
Early Stopping in Contextual Bandits and Inferences
von: Cui, Zihan
Veröffentlicht: (2025)
von: Cui, Zihan
Veröffentlicht: (2025)
Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks
von: Kurumadani, Yuki
Veröffentlicht: (2026)
von: Kurumadani, Yuki
Veröffentlicht: (2026)
Optimal Batched Linear Bandits
von: Ren, Xuanfei, et al.
Veröffentlicht: (2024)
von: Ren, Xuanfei, et al.
Veröffentlicht: (2024)
Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach
von: Reboul, Sebastian, et al.
Veröffentlicht: (2025)
von: Reboul, Sebastian, et al.
Veröffentlicht: (2025)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
von: Agrawal, Shubhada, et al.
Veröffentlicht: (2025)
von: Agrawal, Shubhada, et al.
Veröffentlicht: (2025)
Multitask Learning and Bandits via Robust Statistics
von: Xu, Kan, et al.
Veröffentlicht: (2021)
von: Xu, Kan, et al.
Veröffentlicht: (2021)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
von: Zhao, Qingyue, et al.
Veröffentlicht: (2025)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
von: Yan, Shunxing, et al.
Veröffentlicht: (2026)
von: Yan, Shunxing, et al.
Veröffentlicht: (2026)
The Fragility of Optimized Bandit Algorithms
von: Fan, Lin, et al.
Veröffentlicht: (2021)
von: Fan, Lin, et al.
Veröffentlicht: (2021)
Counterfactual Cocycles: A Framework for Robust and Coherent Counterfactual Transports
von: Dance, Hugh, et al.
Veröffentlicht: (2024)
von: Dance, Hugh, et al.
Veröffentlicht: (2024)
Adaptive Smooth Non-Stationary Bandits
von: Suk, Joe
Veröffentlicht: (2024)
von: Suk, Joe
Veröffentlicht: (2024)
Testing the Feasibility of Linear Programs with Bandit Feedback
von: Gangrade, Aditya, et al.
Veröffentlicht: (2024)
von: Gangrade, Aditya, et al.
Veröffentlicht: (2024)
Truncated LinUCB for Stochastic Linear Bandits
von: Song, Yanglei, et al.
Veröffentlicht: (2022)
von: Song, Yanglei, et al.
Veröffentlicht: (2022)
Design Experiments to Compare Multi-armed Bandit Algorithms
von: Meng, Huiling, et al.
Veröffentlicht: (2026)
von: Meng, Huiling, et al.
Veröffentlicht: (2026)
Hierarchical Clustering With Confidence
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Online Clustering of Data Sequences with Bandit Information
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2025)
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2025)
Can Bayesian Neural Networks Make Confident Predictions?
von: Fisher, Katharine, et al.
Veröffentlicht: (2025)
von: Fisher, Katharine, et al.
Veröffentlicht: (2025)
Improving Kernel-Based Nonasymptotic Simultaneous Confidence Bands
von: Csáji, Balázs Csanád, et al.
Veröffentlicht: (2024)
von: Csáji, Balázs Csanád, et al.
Veröffentlicht: (2024)
Optimal Confidence Band for Kernel Gradient Flow Estimator
von: Cheng, Yuqian, et al.
Veröffentlicht: (2026)
von: Cheng, Yuqian, et al.
Veröffentlicht: (2026)
Asymptotically Optimal Problem-Dependent Bandit Policies for Transfer Learning
von: Prevost, Adrien, et al.
Veröffentlicht: (2025)
von: Prevost, Adrien, et al.
Veröffentlicht: (2025)
Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
von: Ji, Wenlong, et al.
Veröffentlicht: (2025)
von: Ji, Wenlong, et al.
Veröffentlicht: (2025)
Distribution-Free Confidence Ellipsoids for Ridge Regression with PAC Bounds
von: Szentpéteri, Szabolcs, et al.
Veröffentlicht: (2026)
von: Szentpéteri, Szabolcs, et al.
Veröffentlicht: (2026)
Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
On the Upper Bounds for the Matrix Spectral Norm
von: Naumov, Alexey, et al.
Veröffentlicht: (2025)
von: Naumov, Alexey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bayesian Design Principles for Frequentist Sequential Learning
von: Xu, Yunbei, et al.
Veröffentlicht: (2023) -
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
von: Chen, Fan, et al.
Veröffentlicht: (2024) -
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025) -
An Upper Confidence Bound Approach to Estimating the Maximum Mean
von: Kun, Zhang, et al.
Veröffentlicht: (2024) -
On Lai's Upper Confidence Bound in Multi-Armed Bandits
von: Ren, Huachen, et al.
Veröffentlicht: (2024)