Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Junjie, Song, Junhao, Han, Xudong, Bi, Ziqian, Wang, Tianyang, Liang, Chia Xin, Song, Xinyuan, Zhang, Yichao, Niu, Qian, Peng, Benji, Chen, Keyu, Liu, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912334488797184
author Yang, Junjie
Song, Junhao
Han, Xudong
Bi, Ziqian
Wang, Tianyang
Liang, Chia Xin
Song, Xinyuan
Zhang, Yichao
Niu, Qian
Peng, Benji
Chen, Keyu
Liu, Ming
author_facet Yang, Junjie
Song, Junhao
Han, Xudong
Bi, Ziqian
Wang, Tianyang
Liang, Chia Xin
Song, Xinyuan
Zhang, Yichao
Niu, Qian
Peng, Benji
Chen, Keyu
Liu, Ming
contents Knowledge distillation (KD) is a technique for transferring knowledge from complex teacher models to simpler student models, significantly enhancing model efficiency and accuracy. It has demonstrated substantial advancements in various applications including image classification, object detection, language modeling, text classification, and sentiment analysis. Recent innovations in KD methods, such as attention-based approaches, block-wise logit distillation, and decoupling distillation, have notably improved student model performance. These techniques focus on stimulus complexity, attention mechanisms, and global information capture to optimize knowledge transfer. In addition, KD has proven effective in compressing large language models while preserving accuracy, reducing computational overhead, and improving inference speed. This survey synthesizes the latest literature, highlighting key findings, contributions, and future directions in knowledge distillation to provide insights for researchers and practitioners on its evolving role in artificial intelligence and machine learning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
Yang, Junjie
Song, Junhao
Han, Xudong
Bi, Ziqian
Wang, Tianyang
Liang, Chia Xin
Song, Xinyuan
Zhang, Yichao
Niu, Qian
Peng, Benji
Chen, Keyu
Liu, Ming
Computation and Language
Machine Learning
Knowledge distillation (KD) is a technique for transferring knowledge from complex teacher models to simpler student models, significantly enhancing model efficiency and accuracy. It has demonstrated substantial advancements in various applications including image classification, object detection, language modeling, text classification, and sentiment analysis. Recent innovations in KD methods, such as attention-based approaches, block-wise logit distillation, and decoupling distillation, have notably improved student model performance. These techniques focus on stimulus complexity, attention mechanisms, and global information capture to optimize knowledge transfer. In addition, KD has proven effective in compressing large language models while preserving accuracy, reducing computational overhead, and improving inference speed. This survey synthesizes the latest literature, highlighting key findings, contributions, and future directions in knowledge distillation to provide insights for researchers and practitioners on its evolving role in artificial intelligence and machine learning.
title Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.13825