HELLINGER-UCB: A novel algorithm for stochastic multi-armed bandit problem and cold start problem in recommender system
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ruibo, Wang, Jiazhou, Mullhaupt, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
von: Han, Qiyang, et al.
Veröffentlicht: (2024)
von: Han, Qiyang, et al.
Veröffentlicht: (2024)
Fast UCB-type algorithms for stochastic bandits with heavy and super heavy symmetric noise
von: Dorn, Yuriy, et al.
Veröffentlicht: (2024)
von: Dorn, Yuriy, et al.
Veröffentlicht: (2024)
Functional multi-armed bandit and the best function identification problems
von: Dorn, Yuriy, et al.
Veröffentlicht: (2025)
von: Dorn, Yuriy, et al.
Veröffentlicht: (2025)
Leveraging priors on distribution functions for multi-arm bandits
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025)
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025)
Trading off rewards and errors in multi-armed bandits
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
Bounding Neyman-Pearson Region with $f$-Divergences
von: Mullhaupt, Andrew, et al.
Veröffentlicht: (2025)
von: Mullhaupt, Andrew, et al.
Veröffentlicht: (2025)
Minimax-optimal trust-aware multi-armed bandits
von: Cai, Changxiao, et al.
Veröffentlicht: (2024)
von: Cai, Changxiao, et al.
Veröffentlicht: (2024)
Information maximization for a broad variety of multi-armed bandit games
von: Barbier-Chebbah, Alex, et al.
Veröffentlicht: (2025)
von: Barbier-Chebbah, Alex, et al.
Veröffentlicht: (2025)
Spectral bandits for smooth graph functions with applications in recommender systems
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Quantum contextual bandits and recommender systems for quantum data
von: Brahmachari, Shrigyan, et al.
Veröffentlicht: (2023)
von: Brahmachari, Shrigyan, et al.
Veröffentlicht: (2023)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
von: Turinici, Gabriel
Veröffentlicht: (2026)
von: Turinici, Gabriel
Veröffentlicht: (2026)
Extended UCB Policies for Multi-armed Bandit Problems
von: Liu, Keqin, et al.
Veröffentlicht: (2011)
von: Liu, Keqin, et al.
Veröffentlicht: (2011)
Efficient learning by implicit exploration in bandit problems with side observations
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
A more efficient method for large-sample model-free feature screening via multi-armed bandits
von: Ouyang, Xiaxue, et al.
Veröffentlicht: (2025)
von: Ouyang, Xiaxue, et al.
Veröffentlicht: (2025)
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
von: Qin, Hao, et al.
Veröffentlicht: (2025)
von: Qin, Hao, et al.
Veröffentlicht: (2025)
Using causal abstractions to accelerate decision-making in complex bandit problems
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
Offline-to-online hyperparameter transfer for stochastic bandits
von: Sharma, Dravyansh, et al.
Veröffentlicht: (2025)
von: Sharma, Dravyansh, et al.
Veröffentlicht: (2025)
Prior-informed optimization of treatment recommendation via bandit algorithms trained on large language model-processed historical records
von: Nessari, Saman, et al.
Veröffentlicht: (2025)
von: Nessari, Saman, et al.
Veröffentlicht: (2025)
A single algorithm for both restless and rested rotting bandits
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
Efficient kernelized bandit algorithms via exploration distributions
von: Hu, Bingshan, et al.
Veröffentlicht: (2025)
von: Hu, Bingshan, et al.
Veröffentlicht: (2025)
A survey on multi-player bandits
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
Solving cold start in news recommendations: a RippleNet-based system for large scale media outlet
von: Radziszewski, Karol, et al.
Veröffentlicht: (2025)
von: Radziszewski, Karol, et al.
Veröffentlicht: (2025)
Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more
von: Kikkawa, Nobuaki, et al.
Veröffentlicht: (2024)
von: Kikkawa, Nobuaki, et al.
Veröffentlicht: (2024)
Extreme bandits
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Adaptive political surveys and GPT-4: Tackling the cold start problem with simulated user interactions
von: Bachmann, Fynn, et al.
Veröffentlicht: (2025)
von: Bachmann, Fynn, et al.
Veröffentlicht: (2025)
Solving multi-armed bandit problems using a chaotic microresonator comb
von: Cuevas, Jonathan, et al.
Veröffentlicht: (2023)
von: Cuevas, Jonathan, et al.
Veröffentlicht: (2023)
Spectral bandits
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
On the Suboptimality of GP-UCB under Polynomial Effective Optimism
von: Wang, Wenjia, et al.
Veröffentlicht: (2023)
von: Wang, Wenjia, et al.
Veröffentlicht: (2023)
Replicable Bandits with UCB based Exploration
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
Heuristic algorithms for the stochastic critical node detection problem
von: Bayarsaikhan, Tuguldur, et al.
Veröffentlicht: (2025)
von: Bayarsaikhan, Tuguldur, et al.
Veröffentlicht: (2025)
On the optimal regret of collaborative personalized linear bandits
von: Huang, Bruce, et al.
Veröffentlicht: (2025)
von: Huang, Bruce, et al.
Veröffentlicht: (2025)
Active clustering with bandit feedback
von: Thuot, Victor, et al.
Veröffentlicht: (2024)
von: Thuot, Victor, et al.
Veröffentlicht: (2024)
A characterization of sample adaptivity in UCB data
von: Chen, Yilun, et al.
Veröffentlicht: (2025)
von: Chen, Yilun, et al.
Veröffentlicht: (2025)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
von: Qiu, Shuang, et al.
Veröffentlicht: (2022)
von: Qiu, Shuang, et al.
Veröffentlicht: (2022)
Performance-bounded Online Ensemble Learning Method Based on Multi-armed bandits and Its Applications in Real-time Safety Assessment
von: Hu, Songqiao, et al.
Veröffentlicht: (2025)
von: Hu, Songqiao, et al.
Veröffentlicht: (2025)
An accelerated first-order regularized momentum descent ascent algorithm for stochastic nonconvex-concave minimax problems
von: Zhang, Huiling, et al.
Veröffentlicht: (2023)
von: Zhang, Huiling, et al.
Veröffentlicht: (2023)
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
von: Huang, Ruiquan, et al.
Veröffentlicht: (2023)
von: Huang, Ruiquan, et al.
Veröffentlicht: (2023)
A practical PINN framework for multi-scale problems with multi-magnitude loss terms
von: Wang, Yong, et al.
Veröffentlicht: (2023)
von: Wang, Yong, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
von: Han, Qiyang, et al.
Veröffentlicht: (2024) -
Fast UCB-type algorithms for stochastic bandits with heavy and super heavy symmetric noise
von: Dorn, Yuriy, et al.
Veröffentlicht: (2024) -
Functional multi-armed bandit and the best function identification problems
von: Dorn, Yuriy, et al.
Veröffentlicht: (2025) -
Leveraging priors on distribution functions for multi-arm bandits
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025) -
Trading off rewards and errors in multi-armed bandits
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)