Saved in:
Bibliographic Details
Main Author: Ku, Eugene
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.10807
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910369110294528
author Ku, Eugene
author_facet Ku, Eugene
contents Knowledge Distillation (KD) aims to transfer a more capable teacher model's knowledge to a lighter student model in order to improve the efficiency of the model, making it faster and more deployable. However, the student model's optimization process over the noisy pseudo labels (generated by the teacher model) is tricky and the amount of pseudo labels one can generate is limited due to Out of Memory (OOM) error. In this paper, we propose FlyKD (Knowledge Distillation on the Fly) which enables the generation of virtually unlimited number of pseudo labels, coupled with Curriculum Learning that greatly alleviates the optimization process over the noisy pseudo labels. Empirically, we observe that FlyKD outperforms vanilla KD and the renown Local Structure Preserving Graph Convolutional Network (LSPGCN). Lastly, with the success of Curriculum Learning, we shed light on a new research direction of improving optimization over noisy pseudo labels.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10807
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FlyKD: Graph Knowledge Distillation on the Fly with Curriculum Learning
Ku, Eugene
Machine Learning
Knowledge Distillation (KD) aims to transfer a more capable teacher model's knowledge to a lighter student model in order to improve the efficiency of the model, making it faster and more deployable. However, the student model's optimization process over the noisy pseudo labels (generated by the teacher model) is tricky and the amount of pseudo labels one can generate is limited due to Out of Memory (OOM) error. In this paper, we propose FlyKD (Knowledge Distillation on the Fly) which enables the generation of virtually unlimited number of pseudo labels, coupled with Curriculum Learning that greatly alleviates the optimization process over the noisy pseudo labels. Empirically, we observe that FlyKD outperforms vanilla KD and the renown Local Structure Preserving Graph Convolutional Network (LSPGCN). Lastly, with the success of Curriculum Learning, we shed light on a new research direction of improving optimization over noisy pseudo labels.
title FlyKD: Graph Knowledge Distillation on the Fly with Curriculum Learning
topic Machine Learning
url https://arxiv.org/abs/2403.10807