Diversity-Preserving K-Armed Bandits, Revisited

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hadiji, Hédi, Gerchinovitz, Sébastien, Loubes, Jean-Michel, Stoltz, Gilles
Formato: Preprint
Publicado: 2020
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911965203398656
author Hadiji, Hédi
Gerchinovitz, Sébastien
Loubes, Jean-Michel
Stoltz, Gilles
author_facet Hadiji, Hédi
Gerchinovitz, Sébastien
Loubes, Jean-Michel
Stoltz, Gilles
contents We consider the bandit-based framework for diversity-preserving recommendations introduced by Celis et al. (2019), who approached it in the case of a polytope mainly by a reduction to the setting of linear bandits. We design a UCB algorithm using the specific structure of the setting and show that it enjoys a bounded distribution-dependent regret in the natural cases when the optimal mixed actions put some probability mass on all actions (i.e., when diversity is desirable). The regret lower bounds provided show that otherwise, at least when the model is mean-unbounded, a $\ln T$ regret is suffered. We also discuss an example beyond the special case of polytopes.
format Preprint
id arxiv_https___arxiv_org_abs_2010_01874
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Diversity-Preserving K-Armed Bandits, Revisited
Hadiji, Hédi
Gerchinovitz, Sébastien
Loubes, Jean-Michel
Stoltz, Gilles
Machine Learning
We consider the bandit-based framework for diversity-preserving recommendations introduced by Celis et al. (2019), who approached it in the case of a polytope mainly by a reduction to the setting of linear bandits. We design a UCB algorithm using the specific structure of the setting and show that it enjoys a bounded distribution-dependent regret in the natural cases when the optimal mixed actions put some probability mass on all actions (i.e., when diversity is desirable). The regret lower bounds provided show that otherwise, at least when the model is mean-unbounded, a $\ln T$ regret is suffered. We also discuss an example beyond the special case of polytopes.
title Diversity-Preserving K-Armed Bandits, Revisited
topic Machine Learning
url https://arxiv.org/abs/2010.01874