Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khani, Nikhil, Yang, Shuo, Nath, Aniruddh, Liu, Yang, Abbo, Pendo, Wei, Li, Andrews, Shawn, Kula, Maciej, Kahn, Jarrod, Zhao, Zhe, Hong, Lichan, Chi, Ed
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912002641756160
author Khani, Nikhil
Yang, Shuo
Nath, Aniruddh
Liu, Yang
Abbo, Pendo
Wei, Li
Andrews, Shawn
Kula, Maciej
Kahn, Jarrod
Zhao, Zhe
Hong, Lichan
Chi, Ed
author_facet Khani, Nikhil
Yang, Shuo
Nath, Aniruddh
Liu, Yang
Abbo, Pendo
Wei, Li
Andrews, Shawn
Kula, Maciej
Kahn, Jarrod
Zhao, Zhe
Hong, Lichan
Chi, Ed
contents Knowledge Distillation (KD) is a powerful approach for compressing a large model into a smaller, more efficient model, particularly beneficial for latency-sensitive applications like recommender systems. However, current KD research predominantly focuses on Computer Vision (CV) and NLP tasks, overlooking unique data characteristics and challenges inherent to recommender systems. This paper addresses these overlooked challenges, specifically: (1) mitigating data distribution shifts between teacher and student models, (2) efficiently identifying optimal teacher configurations within time and budgetary constraints, and (3) enabling computationally efficient and rapid sharing of teacher labels to support multiple students. We present a robust KD system developed and rigorously evaluated on multiple large-scale personalized video recommendation systems within Google. Our live experiment results demonstrate significant improvements in student model performance while ensuring consistent and reliable generation of high quality teacher labels from a continuous data stream of data.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14678
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
Khani, Nikhil
Yang, Shuo
Nath, Aniruddh
Liu, Yang
Abbo, Pendo
Wei, Li
Andrews, Shawn
Kula, Maciej
Kahn, Jarrod
Zhao, Zhe
Hong, Lichan
Chi, Ed
Information Retrieval
Artificial Intelligence
Machine Learning
Knowledge Distillation (KD) is a powerful approach for compressing a large model into a smaller, more efficient model, particularly beneficial for latency-sensitive applications like recommender systems. However, current KD research predominantly focuses on Computer Vision (CV) and NLP tasks, overlooking unique data characteristics and challenges inherent to recommender systems. This paper addresses these overlooked challenges, specifically: (1) mitigating data distribution shifts between teacher and student models, (2) efficiently identifying optimal teacher configurations within time and budgetary constraints, and (3) enabling computationally efficient and rapid sharing of teacher labels to support multiple students. We present a robust KD system developed and rigorously evaluated on multiple large-scale personalized video recommendation systems within Google. Our live experiment results demonstrate significant improvements in student model performance while ensuring consistent and reliable generation of high quality teacher labels from a continuous data stream of data.
title Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.14678