Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912334488797184 |
|---|---|
| author | Yang, Junjie Song, Junhao Han, Xudong Bi, Ziqian Wang, Tianyang Liang, Chia Xin Song, Xinyuan Zhang, Yichao Niu, Qian Peng, Benji Chen, Keyu Liu, Ming |
| author_facet | Yang, Junjie Song, Junhao Han, Xudong Bi, Ziqian Wang, Tianyang Liang, Chia Xin Song, Xinyuan Zhang, Yichao Niu, Qian Peng, Benji Chen, Keyu Liu, Ming |
| contents | Knowledge distillation (KD) is a technique for transferring knowledge from complex teacher models to simpler student models, significantly enhancing model efficiency and accuracy. It has demonstrated substantial advancements in various applications including image classification, object detection, language modeling, text classification, and sentiment analysis. Recent innovations in KD methods, such as attention-based approaches, block-wise logit distillation, and decoupling distillation, have notably improved student model performance. These techniques focus on stimulus complexity, attention mechanisms, and global information capture to optimize knowledge transfer. In addition, KD has proven effective in compressing large language models while preserving accuracy, reducing computational overhead, and improving inference speed. This survey synthesizes the latest literature, highlighting key findings, contributions, and future directions in knowledge distillation to provide insights for researchers and practitioners on its evolving role in artificial intelligence and machine learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_13825 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models Yang, Junjie Song, Junhao Han, Xudong Bi, Ziqian Wang, Tianyang Liang, Chia Xin Song, Xinyuan Zhang, Yichao Niu, Qian Peng, Benji Chen, Keyu Liu, Ming Computation and Language Machine Learning Knowledge distillation (KD) is a technique for transferring knowledge from complex teacher models to simpler student models, significantly enhancing model efficiency and accuracy. It has demonstrated substantial advancements in various applications including image classification, object detection, language modeling, text classification, and sentiment analysis. Recent innovations in KD methods, such as attention-based approaches, block-wise logit distillation, and decoupling distillation, have notably improved student model performance. These techniques focus on stimulus complexity, attention mechanisms, and global information capture to optimize knowledge transfer. In addition, KD has proven effective in compressing large language models while preserving accuracy, reducing computational overhead, and improving inference speed. This survey synthesizes the latest literature, highlighting key findings, contributions, and future directions in knowledge distillation to provide insights for researchers and practitioners on its evolving role in artificial intelligence and machine learning. |
| title | Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2504.13825 |