Trading off rewards and errors in multi-armed bandits
Fuente:
arXiv
Guardado en:
| Autores principales: | Erraqabi, Akram, Lazaric, Alessandro, Valko, Michal, Brunskill, Emma, Liu, Yun-En |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A single algorithm for both restless and rested rotting bandits
por: Seznec, Julien, et al.
Publicado: (2026)
por: Seznec, Julien, et al.
Publicado: (2026)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
por: Perrault, Pierre, et al.
Publicado: (2026)
por: Perrault, Pierre, et al.
Publicado: (2026)
Pliable rejection sampling
por: Erraqabi, Akram, et al.
Publicado: (2026)
por: Erraqabi, Akram, et al.
Publicado: (2026)
Large-scale semi-supervised learning with online spectral graph sparsification
por: Calandriello, Daniele, et al.
Publicado: (2026)
por: Calandriello, Daniele, et al.
Publicado: (2026)
Analysis of Nystrom method with sequential ridge leverage scores
por: Calandriello, Daniele, et al.
Publicado: (2026)
por: Calandriello, Daniele, et al.
Publicado: (2026)
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
por: Calandriello, Daniele, et al.
Publicado: (2026)
por: Calandriello, Daniele, et al.
Publicado: (2026)
Extreme bandits
por: Carpentier, Alexandra, et al.
Publicado: (2026)
por: Carpentier, Alexandra, et al.
Publicado: (2026)
Revealing graph bandits for maximizing local influence
por: Carpentier, Alexandra, et al.
Publicado: (2026)
por: Carpentier, Alexandra, et al.
Publicado: (2026)
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
por: Audiffren, Julien, et al.
Publicado: (2026)
por: Audiffren, Julien, et al.
Publicado: (2026)
Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
por: Tarbouriech, Jean, et al.
Publicado: (2026)
por: Tarbouriech, Jean, et al.
Publicado: (2026)
Improved large-scale graph learning through ridge spectral sparsification
por: Calandriello, Daniele, et al.
Publicado: (2026)
por: Calandriello, Daniele, et al.
Publicado: (2026)
Spectral bandits for smooth graph functions
por: Valko, Michal, et al.
Publicado: (2026)
por: Valko, Michal, et al.
Publicado: (2026)
Spectral bandits
por: Kocák, Tomáš, et al.
Publicado: (2026)
por: Kocák, Tomáš, et al.
Publicado: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
por: Kocak, Tomas, et al.
Publicado: (2026)
por: Kocak, Tomas, et al.
Publicado: (2026)
Leveraging priors on distribution functions for multi-arm bandits
por: Vashishtha, Sumit, et al.
Publicado: (2025)
por: Vashishtha, Sumit, et al.
Publicado: (2025)
Spectral bandits for smooth graph functions with applications in recommender systems
por: Kocák, Tomáš, et al.
Publicado: (2026)
por: Kocák, Tomáš, et al.
Publicado: (2026)
Minimax-optimal trust-aware multi-armed bandits
por: Cai, Changxiao, et al.
Publicado: (2024)
por: Cai, Changxiao, et al.
Publicado: (2024)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
por: Fiegel, Come, et al.
Publicado: (2026)
por: Fiegel, Come, et al.
Publicado: (2026)
Scale-free adaptive planning for deterministic dynamics & discounted rewards
por: Bartlett, Peter L., et al.
Publicado: (2026)
por: Bartlett, Peter L., et al.
Publicado: (2026)
Functional multi-armed bandit and the best function identification problems
por: Dorn, Yuriy, et al.
Publicado: (2025)
por: Dorn, Yuriy, et al.
Publicado: (2025)
Information maximization for a broad variety of multi-armed bandit games
por: Barbier-Chebbah, Alex, et al.
Publicado: (2025)
por: Barbier-Chebbah, Alex, et al.
Publicado: (2025)
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
por: Han, Qiyang, et al.
Publicado: (2024)
por: Han, Qiyang, et al.
Publicado: (2024)
Leveraging heterogeneous spillover in maximizing contextual bandit rewards
por: Faruk, Ahmed Sayeed, et al.
Publicado: (2023)
por: Faruk, Ahmed Sayeed, et al.
Publicado: (2023)
Bandits on graphs and structures
por: Valko, Michal
Publicado: (2026)
por: Valko, Michal
Publicado: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
por: Valko, Michal
Publicado: (2026)
por: Valko, Michal
Publicado: (2026)
Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more
por: Kikkawa, Nobuaki, et al.
Publicado: (2024)
por: Kikkawa, Nobuaki, et al.
Publicado: (2024)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
por: Turinici, Gabriel
Publicado: (2026)
por: Turinici, Gabriel
Publicado: (2026)
Adaptive multi-fidelity optimization with fast learning rates
por: Fiegel, Come, et al.
Publicado: (2026)
por: Fiegel, Come, et al.
Publicado: (2026)
Best of both worlds: Stochastic & adversarial best-arm identification
por: Abbasi-Yadkori, Yasin, et al.
Publicado: (2026)
por: Abbasi-Yadkori, Yasin, et al.
Publicado: (2026)
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
por: Qin, Hao, et al.
Publicado: (2025)
por: Qin, Hao, et al.
Publicado: (2025)
A more efficient method for large-sample model-free feature screening via multi-armed bandits
por: Ouyang, Xiaxue, et al.
Publicado: (2025)
por: Ouyang, Xiaxue, et al.
Publicado: (2025)
HELLINGER-UCB: A novel algorithm for stochastic multi-armed bandit problem and cold start problem in recommender system
por: Yang, Ruibo, et al.
Publicado: (2024)
por: Yang, Ruibo, et al.
Publicado: (2024)
Minimum mean-squared error estimation with bandit feedback
por: Ghosh, Ayon, et al.
Publicado: (2022)
por: Ghosh, Ayon, et al.
Publicado: (2022)
Feature importance analysis for patient management decisions
por: Valko, Michal, et al.
Publicado: (2026)
por: Valko, Michal, et al.
Publicado: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
por: Neu, Gergely, et al.
Publicado: (2026)
por: Neu, Gergely, et al.
Publicado: (2026)
Distance metric learning for conditional anomaly detection
por: Valko, Michal, et al.
Publicado: (2026)
por: Valko, Michal, et al.
Publicado: (2026)
Learning from a single labeled face and a stream of unlabeled data
por: Kveton, Branislav, et al.
Publicado: (2026)
por: Kveton, Branislav, et al.
Publicado: (2026)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
Active Learning for Stochastic Contextual Linear Bandits
por: Brunskill, Emma, et al.
Publicado: (2026)
por: Brunskill, Emma, et al.
Publicado: (2026)
Reinforcement Learning with Options and State Representation
por: Ghriss, Ayoub, et al.
Publicado: (2024)
por: Ghriss, Ayoub, et al.
Publicado: (2024)
Ejemplares similares
-
A single algorithm for both restless and rested rotting bandits
por: Seznec, Julien, et al.
Publicado: (2026) -
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
por: Perrault, Pierre, et al.
Publicado: (2026) -
Pliable rejection sampling
por: Erraqabi, Akram, et al.
Publicado: (2026) -
Large-scale semi-supervised learning with online spectral graph sparsification
por: Calandriello, Daniele, et al.
Publicado: (2026) -
Analysis of Nystrom method with sequential ridge leverage scores
por: Calandriello, Daniele, et al.
Publicado: (2026)