PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lim, Hyemin, Lee, Jaeyeon, Choi, Dong-Wan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929700663721984
author Lim, Hyemin
Lee, Jaeyeon
Choi, Dong-Wan
author_facet Lim, Hyemin
Lee, Jaeyeon
Choi, Dong-Wan
contents Large pretrained language models such as BERT suffer from slow inference and high memory usage, due to their huge size. Recent approaches to compressing BERT rely on iterative pruning and knowledge distillation, which, however, are often too complicated and computationally intensive. This paper proposes a novel semi-structured one-shot pruning method for BERT, called $\textit{Permutation and Grouping for BERT}$ (PGB), which achieves high compression efficiency and sparsity while preserving accuracy. To this end, PGB identifies important groups of individual weights by permutation and prunes all other weights as a structure in both multi-head attention and feed-forward layers. Furthermore, if no important group is formed in a particular layer, PGB drops the entire layer to produce an even more compact model. Our experimental results on BERT$_{\text{BASE}}$ demonstrate that PGB outperforms the state-of-the-art structured pruning methods in terms of computational cost and accuracy preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation
Lim, Hyemin
Lee, Jaeyeon
Choi, Dong-Wan
Computation and Language
Artificial Intelligence
Large pretrained language models such as BERT suffer from slow inference and high memory usage, due to their huge size. Recent approaches to compressing BERT rely on iterative pruning and knowledge distillation, which, however, are often too complicated and computationally intensive. This paper proposes a novel semi-structured one-shot pruning method for BERT, called $\textit{Permutation and Grouping for BERT}$ (PGB), which achieves high compression efficiency and sparsity while preserving accuracy. To this end, PGB identifies important groups of individual weights by permutation and prunes all other weights as a structure in both multi-head attention and feed-forward layers. Furthermore, if no important group is formed in a particular layer, PGB drops the entire layer to produce an even more compact model. Our experimental results on BERT$_{\text{BASE}}$ demonstrate that PGB outperforms the state-of-the-art structured pruning methods in terms of computational cost and accuracy preservation.
title PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.03984