Offline Local Search for Online Stochastic Bandits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Benadè, Gerdus, Das, Rathish, Lavastida, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How RLHF Amplifies Sycophancy
von: Shapira, Itai, et al.
Veröffentlicht: (2026)
von: Shapira, Itai, et al.
Veröffentlicht: (2026)
Unlearning Offline Stochastic Multi-Armed Bandits
von: Ye, Zichun, et al.
Veröffentlicht: (2026)
von: Ye, Zichun, et al.
Veröffentlicht: (2026)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
von: Han, Zean, et al.
Veröffentlicht: (2026)
von: Han, Zean, et al.
Veröffentlicht: (2026)
Online Bandit Learning with Offline Preference Data for Improved RLHF
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
von: Cheung, Wang Chi, et al.
Veröffentlicht: (2024)
von: Cheung, Wang Chi, et al.
Veröffentlicht: (2024)
Binary Search with Distributional Predictions
von: Dinitz, Michael, et al.
Veröffentlicht: (2024)
von: Dinitz, Michael, et al.
Veröffentlicht: (2024)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
von: Ge, Haosen, et al.
Veröffentlicht: (2024)
von: Ge, Haosen, et al.
Veröffentlicht: (2024)
From Stream to Pool: Pricing Under the Law of Diminishing Marginal Utility
von: Cui, Titing, et al.
Veröffentlicht: (2023)
von: Cui, Titing, et al.
Veröffentlicht: (2023)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
Group-Sensitive Offline Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
Stochastic Online Instrumental Variable Regression: Regrets for Endogeneity and Bandit Feedback
von: Della Vecchia, Riccardo, et al.
Veröffentlicht: (2023)
von: Della Vecchia, Riccardo, et al.
Veröffentlicht: (2023)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
Lipschitz Bandits with Stochastic Delayed Feedback
von: Liu, Zhongxuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxuan, et al.
Veröffentlicht: (2025)
Offline Contextual Bandits in the Presence of New Actions
von: Kishimoto, Ren, et al.
Veröffentlicht: (2026)
von: Kishimoto, Ren, et al.
Veröffentlicht: (2026)
Refined PAC-Bayes Bounds for Offline Bandits
von: Gouverneur, Amaury, et al.
Veröffentlicht: (2025)
von: Gouverneur, Amaury, et al.
Veröffentlicht: (2025)
Offline Learning for Combinatorial Multi-armed Bandits
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
Stochastic Low-rank Tensor Bandits for Multi-dimensional Online Decision Making
von: Zhou, Jie, et al.
Veröffentlicht: (2020)
von: Zhou, Jie, et al.
Veröffentlicht: (2020)
Efficient Clustering in Stochastic Bandits
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2026)
von: Chandran, G Dhinesh, et al.
Veröffentlicht: (2026)
Stochastic Bandits for Egalitarian Assignment
von: Lim, Eugene, et al.
Veröffentlicht: (2024)
von: Lim, Eugene, et al.
Veröffentlicht: (2024)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits
von: Huang, Ziyi, et al.
Veröffentlicht: (2024)
von: Huang, Ziyi, et al.
Veröffentlicht: (2024)
Are Stochastic Multi-objective Bandits Harder than Single-objective Bandits?
von: Guan, Changkun, et al.
Veröffentlicht: (2026)
von: Guan, Changkun, et al.
Veröffentlicht: (2026)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
Leveraging Offline Data in Linear Latent Contextual Bandits
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
Offline Clustering of Linear Bandits: The Power of Clusters under Limited Data
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Batched Stochastic Bandit for Nondegenerate Functions
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Stochastic Bandits Robust to Adversarial Attacks
von: Wang, Xuchuang, et al.
Veröffentlicht: (2024)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2024)
Online Continuous Hyperparameter Optimization for Generalized Linear Contextual Bandits
von: Kang, Yue, et al.
Veröffentlicht: (2023)
von: Kang, Yue, et al.
Veröffentlicht: (2023)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
von: Avin, Chen, et al.
Veröffentlicht: (2025)
von: Avin, Chen, et al.
Veröffentlicht: (2025)
Demystifying Online Clustering of Bandits: Enhanced Exploration Under Stochastic and Smoothed Adversarial Contexts
von: Li, Zhuohua, et al.
Veröffentlicht: (2025)
von: Li, Zhuohua, et al.
Veröffentlicht: (2025)
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
Online Pre-Training for Offline-to-Online Reinforcement Learning
von: Shin, Yongjae, et al.
Veröffentlicht: (2025)
von: Shin, Yongjae, et al.
Veröffentlicht: (2025)
Stochastic $k$-Submodular Bandits with Full Bandit Feedback
von: Nie, Guanyu, et al.
Veröffentlicht: (2024)
von: Nie, Guanyu, et al.
Veröffentlicht: (2024)
Conformal-Style Quantile Analyses for Stochastic Bandits
von: Du, Chengyu, et al.
Veröffentlicht: (2026)
von: Du, Chengyu, et al.
Veröffentlicht: (2026)
Active Learning for Stochastic Contextual Linear Bandits
von: Brunskill, Emma, et al.
Veröffentlicht: (2026)
von: Brunskill, Emma, et al.
Veröffentlicht: (2026)
Stochastic Graph Bandit Learning with Side-Observations
von: Gong, Xueping, et al.
Veröffentlicht: (2023)
von: Gong, Xueping, et al.
Veröffentlicht: (2023)
Stochastic Matching Bandits with Rare Optimization Updates
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
Biased Dueling Bandits with Stochastic Delayed Feedback
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How RLHF Amplifies Sycophancy
von: Shapira, Itai, et al.
Veröffentlicht: (2026) -
Unlearning Offline Stochastic Multi-Armed Bandits
von: Ye, Zichun, et al.
Veröffentlicht: (2026) -
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
von: Han, Zean, et al.
Veröffentlicht: (2026) -
Online Bandit Learning with Offline Preference Data for Improved RLHF
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024) -
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
von: Cheung, Wang Chi, et al.
Veröffentlicht: (2024)