Saved in:
Bibliographic Details
Main Authors: Namkoong, Hongseok, Daulton, Samuel, Bakshy, Eytan
Format: Preprint
Published: 2020
Subjects:
Online Access:https://arxiv.org/abs/2011.14266
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914880954564608
author Namkoong, Hongseok
Daulton, Samuel
Bakshy, Eytan
author_facet Namkoong, Hongseok
Daulton, Samuel
Bakshy, Eytan
contents Thompson sampling (TS) has emerged as a robust technique for contextual bandit problems. However, TS requires posterior inference and optimization for action generation, prohibiting its use in many online platforms where latency and ease of deployment are of concern. We operationalize TS by proposing a novel imitation-learning-based algorithm that distills a TS policy into an explicit policy representation, allowing fast decision-making and easy deployment in mobile and server-based environments. Using batched data collected under the imitation policy, our algorithm iteratively performs offline updates to the TS policy, and learns a new explicit policy representation to imitate it. Empirically, our imitation policy achieves performance comparable to batch TS while allowing more than an order of magnitude reduction in decision-time latency. Buoyed by low latency and simplicity of implementation, our algorithm has been successfully deployed in multiple video upload systems for Meta. Using a randomized controlled trial, we show our algorithm resulted in significant improvements in video quality and watch time.
format Preprint
id arxiv_https___arxiv_org_abs_2011_14266
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
Namkoong, Hongseok
Daulton, Samuel
Bakshy, Eytan
Machine Learning
Artificial Intelligence
Thompson sampling (TS) has emerged as a robust technique for contextual bandit problems. However, TS requires posterior inference and optimization for action generation, prohibiting its use in many online platforms where latency and ease of deployment are of concern. We operationalize TS by proposing a novel imitation-learning-based algorithm that distills a TS policy into an explicit policy representation, allowing fast decision-making and easy deployment in mobile and server-based environments. Using batched data collected under the imitation policy, our algorithm iteratively performs offline updates to the TS policy, and learns a new explicit policy representation to imitate it. Empirically, our imitation policy achieves performance comparable to batch TS while allowing more than an order of magnitude reduction in decision-time latency. Buoyed by low latency and simplicity of implementation, our algorithm has been successfully deployed in multiple video upload systems for Meta. Using a randomized controlled trial, we show our algorithm resulted in significant improvements in video quality and watch time.
title Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2011.14266