Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kolawole, Steven, Dery, Lucio, Kagy, Jean-François, Smith, Virginia, Neubig, Graham, Talwalkar, Ameet
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2402.05406
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908781912260608
author Kolawole, Steven
Dery, Lucio
Kagy, Jean-François
Smith, Virginia
Neubig, Graham
Talwalkar, Ameet
author_facet Kolawole, Steven
Dery, Lucio
Kagy, Jean-François
Smith, Virginia
Neubig, Graham
Talwalkar, Ameet
contents Structured pruning is a promising approach to create smaller, faster large language models. However, existing methods typically rely on computing the gradient via backward passes, which can inflate memory requirements and compute costs. In this work we introduce Bonsai, a gradient-free structured pruning method that eliminates the need for backpropagation, significantly reducing memory requirements and compute costs while achieving state-of-the-art pruning performance. Bonsai uses forward-pass-only perturbative pruning to enable efficient compression of large models on a broader range of hardware configurations. Unlike existing structured pruning approaches, Bonsai not only achieves better compression with fewer resources but also produces models that are twice as fast as those generated by semi-structured pruning. As a concrete demonstration, we use Bonsai to prune 7B and 8B models to 50% sparsity on a single A6000 GPU -- a task challenging for backprop-based methods in memory-constrained settings, as they require 2-3x the memory. Our results show that removing backprop as a requirement not only enables pruning larger models on constrained hardware but can also lead to state-of-the-art efficiency and performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05406
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
Kolawole, Steven
Dery, Lucio
Kagy, Jean-François
Smith, Virginia
Neubig, Graham
Talwalkar, Ameet
Machine Learning
Computation and Language
Structured pruning is a promising approach to create smaller, faster large language models. However, existing methods typically rely on computing the gradient via backward passes, which can inflate memory requirements and compute costs. In this work we introduce Bonsai, a gradient-free structured pruning method that eliminates the need for backpropagation, significantly reducing memory requirements and compute costs while achieving state-of-the-art pruning performance. Bonsai uses forward-pass-only perturbative pruning to enable efficient compression of large models on a broader range of hardware configurations. Unlike existing structured pruning approaches, Bonsai not only achieves better compression with fewer resources but also produces models that are twice as fast as those generated by semi-structured pruning. As a concrete demonstration, we use Bonsai to prune 7B and 8B models to 50% sparsity on a single A6000 GPU -- a task challenging for backprop-based methods in memory-constrained settings, as they require 2-3x the memory. Our results show that removing backprop as a requirement not only enables pruning larger models on constrained hardware but can also lead to state-of-the-art efficiency and performance.
title Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2402.05406