PIP: Perturbation-based Iterative Pruning for Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cao, Yi, Xu, Wei-Jie, Shen, Yucheng, Shi, Weijie, Chan, Chi-Min, Qu, Jianfeng, Xu, Jiajie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911267173695488
author Cao, Yi
Xu, Wei-Jie
Shen, Yucheng
Shi, Weijie
Chan, Chi-Min
Qu, Jianfeng
Xu, Jiajie
author_facet Cao, Yi
Xu, Wei-Jie
Shen, Yucheng
Shi, Weijie
Chan, Chi-Min
Qu, Jianfeng
Xu, Jiajie
contents The rapid increase in the parameter counts of Large Language Models (LLMs), which often reach into the billions or even trillions, presents significant challenges for their practical deployment, particularly in resource-constrained environments. To address this issue, we propose PIP (Perturbation-based Iterative Pruning), a novel double-view structured pruning method to optimize LLMs, which combines information from two different views: the unperturbed view and the perturbed view. With the calculation of gradient differences, PIP iteratively prunes those that struggle to distinguish between these two views. Our experiments show that PIP reduces the parameter count by approximately 20% while retaining over 85% of the original model's accuracy across varied benchmarks. In some cases, the performance of the pruned model is within 5% of the unpruned version, demonstrating PIP's ability to preserve key aspects of model effectiveness. Moreover, PIP consistently outperforms existing state-of-the-art (SOTA) structured pruning methods, establishing it as a leading technique for optimizing LLMs in constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIP: Perturbation-based Iterative Pruning for Large Language Models
Cao, Yi
Xu, Wei-Jie
Shen, Yucheng
Shi, Weijie
Chan, Chi-Min
Qu, Jianfeng
Xu, Jiajie
Machine Learning
Computation and Language
The rapid increase in the parameter counts of Large Language Models (LLMs), which often reach into the billions or even trillions, presents significant challenges for their practical deployment, particularly in resource-constrained environments. To address this issue, we propose PIP (Perturbation-based Iterative Pruning), a novel double-view structured pruning method to optimize LLMs, which combines information from two different views: the unperturbed view and the perturbed view. With the calculation of gradient differences, PIP iteratively prunes those that struggle to distinguish between these two views. Our experiments show that PIP reduces the parameter count by approximately 20% while retaining over 85% of the original model's accuracy across varied benchmarks. In some cases, the performance of the pruned model is within 5% of the unpruned version, demonstrating PIP's ability to preserve key aspects of model effectiveness. Moreover, PIP consistently outperforms existing state-of-the-art (SOTA) structured pruning methods, establishing it as a leading technique for optimizing LLMs in constrained environments.
title PIP: Perturbation-based Iterative Pruning for Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2501.15278