Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ryu, J. Jon, Kwon, Jeongyeol, Koppe, Benjamin, Jun, Kwang-Sung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
von: Li, Yinan, et al.
Veröffentlicht: (2025)
von: Li, Yinan, et al.
Veröffentlicht: (2025)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
Lower Bounds for Time-Varying Kernelized Bandits
von: Cai, Xu, et al.
Veröffentlicht: (2024)
von: Cai, Xu, et al.
Veröffentlicht: (2024)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
Regret Bounds for Noise-Free Cascaded Kernelized Bandits
von: Li, Zihan, et al.
Veröffentlicht: (2022)
von: Li, Zihan, et al.
Veröffentlicht: (2022)
Quantum-Enhanced Neural Contextual Bandit Algorithms
von: Huang, Yuqi, et al.
Veröffentlicht: (2026)
von: Huang, Yuqi, et al.
Veröffentlicht: (2026)
Harnessing the Power of Federated Learning in Federated Contextual Bandits
von: Shi, Chengshuai, et al.
Veröffentlicht: (2023)
von: Shi, Chengshuai, et al.
Veröffentlicht: (2023)
Beam-aware Kernelized Contextual Bandits for User Association and Beamforming in mmWave Vehicular Networks
von: He, Xiaoyang, et al.
Veröffentlicht: (2026)
von: He, Xiaoyang, et al.
Veröffentlicht: (2026)
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)
von: Pacchiano, Aldo
Veröffentlicht: (2024)
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
von: Chen, Fan, et al.
Veröffentlicht: (2024)
von: Chen, Fan, et al.
Veröffentlicht: (2024)
Contrastive Predictive Coding Done Right for Mutual Information Estimation
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
Comparing Comparators in Generalization Bounds
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
Minimax Optimal Algorithms with Fixed-$k$-Nearest Neighbors
von: Ryu, J. Jon, et al.
Veröffentlicht: (2022)
von: Ryu, J. Jon, et al.
Veröffentlicht: (2022)
On the Complexity of First-Order Methods in Stochastic Bilevel Optimization
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Restless Linear Bandits
von: Khaleghi, Azadeh
Veröffentlicht: (2024)
von: Khaleghi, Azadeh
Veröffentlicht: (2024)
Order Optimal Regret Bounds for Sharpe Ratio Optimization under Thompson Sampling
von: Shah, Mohammad Taha, et al.
Veröffentlicht: (2025)
von: Shah, Mohammad Taha, et al.
Veröffentlicht: (2025)
A Classification View on Meta Learning Bandits
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
Shape of Memory: a Geometric Analysis of Machine Unlearning in Second-Order Optimizers
von: Stewart, Kennon
Veröffentlicht: (2026)
von: Stewart, Kennon
Veröffentlicht: (2026)
Optimal Clustering with Bandit Feedback
von: Yang, Junwen, et al.
Veröffentlicht: (2022)
von: Yang, Junwen, et al.
Veröffentlicht: (2022)
Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning
von: Barnes, L. P., et al.
Veröffentlicht: (2022)
von: Barnes, L. P., et al.
Veröffentlicht: (2022)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
von: Qin, Hao, et al.
Veröffentlicht: (2023)
von: Qin, Hao, et al.
Veröffentlicht: (2023)
Batched Kernelized Bandits: Refinements and Extensions
von: Ma, Chenkai, et al.
Veröffentlicht: (2026)
von: Ma, Chenkai, et al.
Veröffentlicht: (2026)
Optimal Arm Elimination Algorithms for Combinatorial Bandits
von: Wen, Yuxiao, et al.
Veröffentlicht: (2025)
von: Wen, Yuxiao, et al.
Veröffentlicht: (2025)
Bandit Convex Optimization with Gradient Prediction Adaptivity
von: Wang, Shuche, et al.
Veröffentlicht: (2026)
von: Wang, Shuche, et al.
Veröffentlicht: (2026)
Conversational Dueling Bandits in Generalized Linear Models
von: Yang, Shuhua, et al.
Veröffentlicht: (2024)
von: Yang, Shuhua, et al.
Veröffentlicht: (2024)
Calibrated Recommendations with Contextual Bandits
von: Feijer, Diego, et al.
Veröffentlicht: (2025)
von: Feijer, Diego, et al.
Veröffentlicht: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
Competing Bandits in Matching Markets via Super Stability
von: Basu, Soumya
Veröffentlicht: (2025)
von: Basu, Soumya
Veröffentlicht: (2025)
Quantile Multi-Armed Bandits with 1-bit Feedback
von: Lau, Ivan, et al.
Veröffentlicht: (2025)
von: Lau, Ivan, et al.
Veröffentlicht: (2025)
Online Clustering of Data Sequences with Bandit Information
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2025)
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2025)
A Modularized Framework for Piecewise-Stationary Restless Bandits
von: Li, Kuan-Ta, et al.
Veröffentlicht: (2026)
von: Li, Kuan-Ta, et al.
Veröffentlicht: (2026)
Improved Regret Bounds for Online Fair Division with Bandit Learning
von: Schiffer, Benjamin, et al.
Veröffentlicht: (2025)
von: Schiffer, Benjamin, et al.
Veröffentlicht: (2025)
Optimal Best Arm Identification with Fixed Confidence in Restless Bandits
von: Karthik, P. N., et al.
Veröffentlicht: (2023)
von: Karthik, P. N., et al.
Veröffentlicht: (2023)
Regret Tail Characterization of Optimal Bandit Algorithms with Generic Rewards
von: Panda, Subhodip, et al.
Veröffentlicht: (2026)
von: Panda, Subhodip, et al.
Veröffentlicht: (2026)
Indexed Minimum Empirical Divergence-Based Algorithms for Linear Bandits
von: Bian, Jie, et al.
Veröffentlicht: (2024)
von: Bian, Jie, et al.
Veröffentlicht: (2024)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
To Switch or Not to Switch? Balanced Policy Switching in Offline Reinforcement Learning
von: Ma, Tao, et al.
Veröffentlicht: (2024)
von: Ma, Tao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
von: Li, Yinan, et al.
Veröffentlicht: (2025) -
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
von: Tajdini, Artin, et al.
Veröffentlicht: (2025) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026) -
Lower Bounds for Time-Varying Kernelized Bandits
von: Cai, Xu, et al.
Veröffentlicht: (2024) -
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)