Spectral bandits for smooth graph functions with applications in recommender systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Kocák, Tomáš, Valko, Michal, Munos, Rémi, Kveton, Branislav, Agrawal, Shipra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spectral bandits for smooth graph functions
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Spectral bandits
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
Spectral Thompson sampling
di: Kocak, Tomas, et al.
Pubblicazione: (2026)
di: Kocak, Tomas, et al.
Pubblicazione: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
di: Kocak, Tomas, et al.
Pubblicazione: (2026)
di: Kocak, Tomas, et al.
Pubblicazione: (2026)
Black-box optimization of noisy functions with unknown smoothness
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Bandits attack function optimization
di: Preux, Philippe, et al.
Pubblicazione: (2026)
di: Preux, Philippe, et al.
Pubblicazione: (2026)
Semi-supervised learning with max-margin graph cuts
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Revealing graph bandits for maximizing local influence
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
Conditional anomaly detection using soft harmonic functions: An application to clinical alerting
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Online learning with Erdős-Rényi side-observation graphs
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
Stochastic simultaneous optimistic optimization
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Extreme bandits
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
Online semi-supervised perception: Real-time learning without explicit feedback
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
Conditional anomaly detection with soft harmonic functions
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
di: Perrault, Pierre, et al.
Pubblicazione: (2026)
di: Perrault, Pierre, et al.
Pubblicazione: (2026)
Online learning with noisy side observations
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
VA-learning as a more efficient alternative to Q-learning
di: Tang, Yunhao, et al.
Pubblicazione: (2023)
di: Tang, Yunhao, et al.
Pubblicazione: (2023)
Evidence-based anomaly detection in clinical domains
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
Bandits on graphs and structures
di: Valko, Michal
Pubblicazione: (2026)
di: Valko, Michal
Pubblicazione: (2026)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
Cross-Validated Off-Policy Evaluation
di: Cief, Matej, et al.
Pubblicazione: (2024)
di: Cief, Matej, et al.
Pubblicazione: (2024)
A single algorithm for both restless and rested rotting bandits
di: Seznec, Julien, et al.
Pubblicazione: (2026)
di: Seznec, Julien, et al.
Pubblicazione: (2026)
Planning in entropy-regularized Markov decision processes and games
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
di: Valko, Michal
Pubblicazione: (2026)
di: Valko, Michal
Pubblicazione: (2026)
Trading off rewards and errors in multi-armed bandits
di: Erraqabi, Akram, et al.
Pubblicazione: (2026)
di: Erraqabi, Akram, et al.
Pubblicazione: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
di: Cief, Matej, et al.
Pubblicazione: (2022)
di: Cief, Matej, et al.
Pubblicazione: (2022)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
di: Fiegel, Come, et al.
Pubblicazione: (2026)
di: Fiegel, Come, et al.
Pubblicazione: (2026)
Optimistic Q-learning for average reward and episodic reinforcement learning
di: Agrawal, Priyank, et al.
Pubblicazione: (2024)
di: Agrawal, Priyank, et al.
Pubblicazione: (2024)
On two ways to use determinantal point processes for Monte Carlo integration
di: Gautier, Guillaume, et al.
Pubblicazione: (2026)
di: Gautier, Guillaume, et al.
Pubblicazione: (2026)
LLM-as-Judge on a Budget
di: Saha, Aadirupa, et al.
Pubblicazione: (2026)
di: Saha, Aadirupa, et al.
Pubblicazione: (2026)
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
Quantum contextual bandits and recommender systems for quantum data
di: Brahmachari, Shrigyan, et al.
Pubblicazione: (2023)
di: Brahmachari, Shrigyan, et al.
Pubblicazione: (2023)
On a few pitfalls in KL divergence gradient estimation for RL
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
Q-learning with Posterior Sampling
di: Agrawal, Priyank, et al.
Pubblicazione: (2025)
di: Agrawal, Priyank, et al.
Pubblicazione: (2025)
Large-scale semi-supervised learning with online spectral graph sparsification
di: Calandriello, Daniele, et al.
Pubblicazione: (2026)
di: Calandriello, Daniele, et al.
Pubblicazione: (2026)
Model-free Posterior Sampling via Learning Rate Randomization
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
Super-Exponential Regret for UCT, AlphaGo and Variants
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
Dynamic Pricing and Learning with Long-term Reference Effects
di: Agrawal, Shipra, et al.
Pubblicazione: (2024)
di: Agrawal, Shipra, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Spectral bandits for smooth graph functions
di: Valko, Michal, et al.
Pubblicazione: (2026) -
Spectral bandits
di: Kocák, Tomáš, et al.
Pubblicazione: (2026) -
Spectral Thompson sampling
di: Kocak, Tomas, et al.
Pubblicazione: (2026) -
Efficient learning by implicit exploration in bandit problems with side observations
di: Kocak, Tomas, et al.
Pubblicazione: (2026) -
Black-box optimization of noisy functions with unknown smoothness
di: Grill, Jean-Bastien, et al.
Pubblicazione: (2026)