Delta Knowledge Distillation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Yihan, Kang, Yanbin, Xing, Zhengming, Jiang, Ruijie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908545374486528
author Cao, Yihan
Kang, Yanbin
Xing, Zhengming
Jiang, Ruijie
author_facet Cao, Yihan
Kang, Yanbin
Xing, Zhengming
Jiang, Ruijie
contents Knowledge distillation (KD) is a widely adopted approach for compressing large neural networks by transferring knowledge from a large teacher model to a smaller student model. In the context of large language models, token level KD, typically minimizing the KL divergence between student output distribution and teacher output distribution, has shown strong empirical performance. However, prior work assumes student output distribution and teacher output distribution share the same optimal representation space, a premise that may not hold in many cases. To solve this problem, we propose Delta Knowledge Distillation (Delta-KD), a novel extension of token level KD that encourages the student to approximate an optimal representation space by explicitly preserving the distributional shift Delta introduced during the teacher's supervised finetuning (SFT). Empirical results on ROUGE metrics demonstrate that Delta KD substantially improves student performance while preserving more of the teacher's knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Delta Knowledge Distillation for Large Language Models
Cao, Yihan
Kang, Yanbin
Xing, Zhengming
Jiang, Ruijie
Computation and Language
Artificial Intelligence
Machine Learning
Knowledge distillation (KD) is a widely adopted approach for compressing large neural networks by transferring knowledge from a large teacher model to a smaller student model. In the context of large language models, token level KD, typically minimizing the KL divergence between student output distribution and teacher output distribution, has shown strong empirical performance. However, prior work assumes student output distribution and teacher output distribution share the same optimal representation space, a premise that may not hold in many cases. To solve this problem, we propose Delta Knowledge Distillation (Delta-KD), a novel extension of token level KD that encourages the student to approximate an optimal representation space by explicitly preserving the distributional shift Delta introduced during the teacher's supervised finetuning (SFT). Empirical results on ROUGE metrics demonstrate that Delta KD substantially improves student performance while preserving more of the teacher's knowledge.
title Delta Knowledge Distillation for Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.14526