Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shuqi, He, Bowei, Wu, Han, Song, Linqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910963774521344
author Liu, Shuqi
He, Bowei
Wu, Han
Song, Linqi
author_facet Liu, Shuqi
He, Bowei
Wu, Han
Song, Linqi
contents Post-training pruning has emerged as a crucial optimization technique as large language models (LLMs) continue to grow rapidly. However, the significant variations in weight distributions across different LLMs make fixed pruning strategies inadequate for multiple models. In this paper, we introduce \textbf{\textsc{OptiShear}}, an efficient evolutionary optimization framework for adaptive LLM pruning. Our framework features two key innovations: an effective search space built on our Meta pruning metric to handle diverse weight distributions, and a model-wise reconstruction error for rapid evaluation during search trials. We employ Non-dominated Sorting Genetic Algorithm III (NSGA-III) to optimize both pruning metrics and layerwise sparsity ratios. Through extensive evaluation on LLaMA-1/2/3 and Mistral models (7B-70B) across multiple benchmarks, we demonstrate that our adaptive pruning metrics consistently outperform existing methods. Additionally, our discovered layerwise sparsity ratios enhance the effectiveness of other pruning metrics. The framework exhibits strong cross-task and cross-model generalizability, providing a cost-effective solution for model compression.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
Liu, Shuqi
He, Bowei
Wu, Han
Song, Linqi
Computation and Language
Post-training pruning has emerged as a crucial optimization technique as large language models (LLMs) continue to grow rapidly. However, the significant variations in weight distributions across different LLMs make fixed pruning strategies inadequate for multiple models. In this paper, we introduce \textbf{\textsc{OptiShear}}, an efficient evolutionary optimization framework for adaptive LLM pruning. Our framework features two key innovations: an effective search space built on our Meta pruning metric to handle diverse weight distributions, and a model-wise reconstruction error for rapid evaluation during search trials. We employ Non-dominated Sorting Genetic Algorithm III (NSGA-III) to optimize both pruning metrics and layerwise sparsity ratios. Through extensive evaluation on LLaMA-1/2/3 and Mistral models (7B-70B) across multiple benchmarks, we demonstrate that our adaptive pruning metrics consistently outperform existing methods. Additionally, our discovered layerwise sparsity ratios enhance the effectiveness of other pruning metrics. The framework exhibits strong cross-task and cross-model generalizability, providing a cost-effective solution for model compression.
title Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2502.10735