Diversity-Preserving K-Armed Bandits, Revisited
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911965203398656 |
|---|---|
| author | Hadiji, Hédi Gerchinovitz, Sébastien Loubes, Jean-Michel Stoltz, Gilles |
| author_facet | Hadiji, Hédi Gerchinovitz, Sébastien Loubes, Jean-Michel Stoltz, Gilles |
| contents | We consider the bandit-based framework for diversity-preserving recommendations introduced by Celis et al. (2019), who approached it in the case of a polytope mainly by a reduction to the setting of linear bandits. We design a UCB algorithm using the specific structure of the setting and show that it enjoys a bounded distribution-dependent regret in the natural cases when the optimal mixed actions put some probability mass on all actions (i.e., when diversity is desirable). The regret lower bounds provided show that otherwise, at least when the model is mean-unbounded, a $\ln T$ regret is suffered. We also discuss an example beyond the special case of polytopes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2010_01874 |
| institution | arXiv |
| publishDate | 2020 |
| record_format | arxiv |
| spellingShingle | Diversity-Preserving K-Armed Bandits, Revisited Hadiji, Hédi Gerchinovitz, Sébastien Loubes, Jean-Michel Stoltz, Gilles Machine Learning We consider the bandit-based framework for diversity-preserving recommendations introduced by Celis et al. (2019), who approached it in the case of a polytope mainly by a reduction to the setting of linear bandits. We design a UCB algorithm using the specific structure of the setting and show that it enjoys a bounded distribution-dependent regret in the natural cases when the optimal mixed actions put some probability mass on all actions (i.e., when diversity is desirable). The regret lower bounds provided show that otherwise, at least when the model is mean-unbounded, a $\ln T$ regret is suffered. We also discuss an example beyond the special case of polytopes. |
| title | Diversity-Preserving K-Armed Bandits, Revisited |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2010.01874 |