LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Runming, Wu, Taiqiang, Wang, Jiahao, Hu, Pengfei, Wu, Yik-Chung, Wong, Ngai, Yang, Yujiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913705303736320
author Yang, Runming
Wu, Taiqiang
Wang, Jiahao
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
author_facet Yang, Runming
Wu, Taiqiang
Wang, Jiahao
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
contents Knowledge distillation (KD) has been a predominant method for compressing Large Language Models (LLMs). In this paper, we first revisit KD and Low-Rank Adaption (LoRA) and demonstrate that they follow the same paradigm. Inspired by this observation, we propose a parameter-efficient knowledge distillation method, LLM-NEO, which integrates LoRA into KD to improve the efficiency of knowledge transfer. After that, we summarize some valuable guidelines for the hyperparameters in LLM-NEO. Experimental results on compressing Llama 2 and Llama 3.2 show that LLM-NEO outperforms various baselines. Further analysis demonstrates the robustness of the proposed LLM-NEO on variants of LoRA. The code and trained models are available at [Github](https://github.com/yang3121099/LLM-Neo).
format Preprint
id arxiv_https___arxiv_org_abs_2411_06839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
Yang, Runming
Wu, Taiqiang
Wang, Jiahao
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
Computation and Language
Artificial Intelligence
Machine Learning
Knowledge distillation (KD) has been a predominant method for compressing Large Language Models (LLMs). In this paper, we first revisit KD and Low-Rank Adaption (LoRA) and demonstrate that they follow the same paradigm. Inspired by this observation, we propose a parameter-efficient knowledge distillation method, LLM-NEO, which integrates LoRA into KD to improve the efficiency of knowledge transfer. After that, we summarize some valuable guidelines for the hyperparameters in LLM-NEO. Experimental results on compressing Llama 2 and Llama 3.2 show that LLM-NEO outperforms various baselines. Further analysis demonstrates the robustness of the proposed LLM-NEO on variants of LoRA. The code and trained models are available at [Github](https://github.com/yang3121099/LLM-Neo).
title LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.06839