Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Anshumann, Zaidi, Mohd Abbas, Kedia, Akhil, Ahn, Jinwoo, Kwon, Taehwak, Lee, Kangwook, Lee, Haejun, Lee, Joohyung
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915407197110272
author Anshumann
Zaidi, Mohd Abbas
Kedia, Akhil
Ahn, Jinwoo
Kwon, Taehwak
Lee, Kangwook
Lee, Haejun
Lee, Joohyung
author_facet Anshumann
Zaidi, Mohd Abbas
Kedia, Akhil
Ahn, Jinwoo
Kwon, Taehwak
Lee, Kangwook
Lee, Haejun
Lee, Joohyung
contents Knowledge distillation can be a cost-effective technique to distill knowledge in Large Language Models, if the teacher output logits can be pre-computed and cached. However, successfully applying this to pre-training remains largely unexplored. In this work, we prove that naive approaches for sparse knowledge distillation such as caching Top-K probabilities, while intuitive, provide biased estimates of teacher probability distribution to the student, resulting in suboptimal performance and calibration. We propose an importance-sampling-based method `Random Sampling Knowledge Distillation', which provides unbiased estimates, preserves the gradient in expectation, and requires storing significantly sparser logits. Our method enables faster training of student models with marginal overhead (<10%) compared to cross-entropy based training, while maintaining competitive performance compared to full distillation, across a range of model sizes from 300M to 3B.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
Anshumann
Zaidi, Mohd Abbas
Kedia, Akhil
Ahn, Jinwoo
Kwon, Taehwak
Lee, Kangwook
Lee, Haejun
Lee, Joohyung
Machine Learning
Artificial Intelligence
Computation and Language
68T50
I.2.7
Knowledge distillation can be a cost-effective technique to distill knowledge in Large Language Models, if the teacher output logits can be pre-computed and cached. However, successfully applying this to pre-training remains largely unexplored. In this work, we prove that naive approaches for sparse knowledge distillation such as caching Top-K probabilities, while intuitive, provide biased estimates of teacher probability distribution to the student, resulting in suboptimal performance and calibration. We propose an importance-sampling-based method `Random Sampling Knowledge Distillation', which provides unbiased estimates, preserves the gradient in expectation, and requires storing significantly sparser logits. Our method enables faster training of student models with marginal overhead (<10%) compared to cross-entropy based training, while maintaining competitive performance compared to full distillation, across a range of model sizes from 300M to 3B.
title Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2503.16870