Reconstruct the Pruned Model without Any Retraining

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Pingjie, Fan, Ziqing, Hu, Shengchao, Chen, Zhe, Wang, Yanfeng, Wang, Yu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916329390342144
author Wang, Pingjie
Fan, Ziqing
Hu, Shengchao
Chen, Zhe
Wang, Yanfeng
Wang, Yu
author_facet Wang, Pingjie
Fan, Ziqing
Hu, Shengchao
Chen, Zhe
Wang, Yanfeng
Wang, Yu
contents Structured pruning is a promising hardware-friendly compression technique for large language models (LLMs), which is expected to be retraining-free to avoid the enormous retraining cost. This retraining-free paradigm involves (1) pruning criteria to define the architecture and (2) distortion reconstruction to restore performance. However, existing methods often emphasize pruning criteria while using reconstruction techniques that are specific to certain modules or criteria, resulting in limited generalizability. To address this, we introduce the Linear Interpolation-based Adaptive Reconstruction (LIAR) framework, which is both efficient and effective. LIAR does not require back-propagation or retraining and is compatible with various pruning criteria and modules. By applying linear interpolation to the preserved weights, LIAR minimizes reconstruction error and effectively reconstructs the pruned output. Our evaluations on benchmarks such as GLUE, SQuAD, WikiText, and common sense reasoning show that LIAR enables a BERT model to maintain 98% accuracy even after removing 50% of its parameters and achieves top performance for LLaMA in just a few minutes.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reconstruct the Pruned Model without Any Retraining
Wang, Pingjie
Fan, Ziqing
Hu, Shengchao
Chen, Zhe
Wang, Yanfeng
Wang, Yu
Machine Learning
Structured pruning is a promising hardware-friendly compression technique for large language models (LLMs), which is expected to be retraining-free to avoid the enormous retraining cost. This retraining-free paradigm involves (1) pruning criteria to define the architecture and (2) distortion reconstruction to restore performance. However, existing methods often emphasize pruning criteria while using reconstruction techniques that are specific to certain modules or criteria, resulting in limited generalizability. To address this, we introduce the Linear Interpolation-based Adaptive Reconstruction (LIAR) framework, which is both efficient and effective. LIAR does not require back-propagation or retraining and is compatible with various pruning criteria and modules. By applying linear interpolation to the preserved weights, LIAR minimizes reconstruction error and effectively reconstructs the pruned output. Our evaluations on benchmarks such as GLUE, SQuAD, WikiText, and common sense reasoning show that LIAR enables a BERT model to maintain 98% accuracy even after removing 50% of its parameters and achieves top performance for LLaMA in just a few minutes.
title Reconstruct the Pruned Model without Any Retraining
topic Machine Learning
url https://arxiv.org/abs/2407.13331