Ranking In Generalized Linear Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914626385477632 |
|---|---|
| author | Shidani, Amitis Deligiannidis, George Doucet, Arnaud |
| author_facet | Shidani, Amitis Deligiannidis, George Doucet, Arnaud |
| contents | We study the ranking problem in generalized linear bandits. At each time, the learning agent selects an ordered list of items and observes stochastic outcomes. In recommendation systems, displaying an ordered list of the most attractive items is not always optimal as both position and item dependencies result in a complex reward function. A very naive example is the lack of diversity when all the most attractive items are from the same category. We model the position and item dependencies in the ordered list and design UCB and Thompson Sampling type algorithms for this problem. Our work generalizes existing studies in several directions, including position dependencies where position discount is a particular case, and connecting the ranking problem to graph theory. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2207_00109 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Ranking In Generalized Linear Bandits Shidani, Amitis Deligiannidis, George Doucet, Arnaud Machine Learning Information Retrieval Optimization and Control We study the ranking problem in generalized linear bandits. At each time, the learning agent selects an ordered list of items and observes stochastic outcomes. In recommendation systems, displaying an ordered list of the most attractive items is not always optimal as both position and item dependencies result in a complex reward function. A very naive example is the lack of diversity when all the most attractive items are from the same category. We model the position and item dependencies in the ordered list and design UCB and Thompson Sampling type algorithms for this problem. Our work generalizes existing studies in several directions, including position dependencies where position discount is a particular case, and connecting the ranking problem to graph theory. |
| title | Ranking In Generalized Linear Bandits |
| topic | Machine Learning Information Retrieval Optimization and Control |
| url | https://arxiv.org/abs/2207.00109 |