Covariance-adapting algorithm for semi-bandits with application to sparse rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Perrault, Pierre, Perchet, Vianney, Valko, Michal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
by: Fiegel, Côme, et al.
Published: (2026)
by: Fiegel, Côme, et al.
Published: (2026)
A single algorithm for both restless and rested rotting bandits
by: Seznec, Julien, et al.
Published: (2026)
by: Seznec, Julien, et al.
Published: (2026)
Extreme bandits
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Revealing graph bandits for maximizing local influence
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Scale-free adaptive planning for deterministic dynamics & discounted rewards
by: Bartlett, Peter L., et al.
Published: (2026)
by: Bartlett, Peter L., et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Budgeted Online Influence Maximization
by: Perrault, Pierre, et al.
Published: (2026)
by: Perrault, Pierre, et al.
Published: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Learning in Prophet Inequalities with Noisy Observations
by: Kim, Jung-hun, et al.
Published: (2026)
by: Kim, Jung-hun, et al.
Published: (2026)
Comparing Uniform Price and Discriminatory Multi-Unit Auctions through Regret Minimization
by: Potfer, Marius, et al.
Published: (2025)
by: Potfer, Marius, et al.
Published: (2025)
Non-clairvoyant Scheduling with Partial Predictions
by: Benomar, Ziyad, et al.
Published: (2024)
by: Benomar, Ziyad, et al.
Published: (2024)
On Tradeoffs in Learning-Augmented Algorithms
by: Benomar, Ziyad, et al.
Published: (2025)
by: Benomar, Ziyad, et al.
Published: (2025)
Online Packet Scheduling with Deadlines and Learning
by: Genalti, Gianmarco, et al.
Published: (2026)
by: Genalti, Gianmarco, et al.
Published: (2026)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Large-scale semi-supervised learning with online spectral graph sparsification
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Leveraging heterogeneous spillover in maximizing contextual bandit rewards
by: Faruk, Ahmed Sayeed, et al.
Published: (2023)
by: Faruk, Ahmed Sayeed, et al.
Published: (2023)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Active multiple matrix completion with adaptive confidence sets
by: Locatelli, Andrea, et al.
Published: (2026)
by: Locatelli, Andrea, et al.
Published: (2026)
Bandits on graphs and structures
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more
by: Kikkawa, Nobuaki, et al.
Published: (2024)
by: Kikkawa, Nobuaki, et al.
Published: (2024)
Demonstration-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Strategic Multi-Armed Bandit Problems Under Debt-Free Reporting
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
Online semi-supervised perception: Real-time learning without explicit feedback
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Adaptive Bandit Algorithms for Contextual Matching Markets
by: Lin, Shiyun, et al.
Published: (2026)
by: Lin, Shiyun, et al.
Published: (2026)
Stable Matching with Ties: Approximation Ratios and Learning
by: Lin, Shiyun, et al.
Published: (2024)
by: Lin, Shiyun, et al.
Published: (2024)
On the Hardness of Reinforcement Learning with Transition Look-Ahead
by: Pla, Corentin, et al.
Published: (2025)
by: Pla, Corentin, et al.
Published: (2025)
Model-free Posterior Sampling via Learning Rate Randomization
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Sharper Perturbed-Kullback-Leibler Exponential Tail Bounds for Beta and Dirichlet Distributions
by: Perrault, Pierre
Published: (2025)
by: Perrault, Pierre
Published: (2025)
Feature importance analysis for patient management decisions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
by: Neu, Gergely, et al.
Published: (2026)
by: Neu, Gergely, et al.
Published: (2026)
Distance metric learning for conditional anomaly detection
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Learning from a single labeled face and a stream of unlabeled data
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Similar Items
-
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026) -
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026) -
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022) -
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026) -
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
by: Fiegel, Côme, et al.
Published: (2026)