Learnable Permutation for Structured Sparsity on Transformer Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zekai, Liu, Ji, Li, Guanchen, Xu, Yixing, Liu, Ziqiong, Yin, Xuanwu, Li, Dong, Barsoum, Emad
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911411224969216
author Li, Zekai
Liu, Ji
Li, Guanchen
Xu, Yixing
Liu, Ziqiong
Yin, Xuanwu
Li, Dong
Barsoum, Emad
author_facet Li, Zekai
Liu, Ji
Li, Guanchen
Xu, Yixing
Liu, Ziqiong
Yin, Xuanwu
Li, Dong
Barsoum, Emad
contents Structured sparsity has emerged as a popular model pruning technique, widely adopted in various architectures, including CNNs, Transformer models, and especially large language models (LLMs) in recent years. A promising direction to further improve post-pruning performance is weight permutation, which reorders model weights into patterns more amenable to pruning. However, the exponential growth of the permutation search space with the scale of Transformer architectures forces most methods to rely on greedy or heuristic algorithms, limiting the effectiveness of reordering. In this work, we propose a novel end-to-end learnable permutation framework. Our method introduces a learnable permutation cost matrix to quantify the cost of swapping any two input channels of a given weight matrix, a differentiable bipartite matching solver to obtain the optimal binary permutation matrix given a cost matrix, and a sparsity optimization loss function to directly optimize the permutation operator. We extensively validate our approach on vision and language Transformers, demonstrating that our method achieves state-of-the-art permutation results for structured sparsity.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22980
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learnable Permutation for Structured Sparsity on Transformer Models
Li, Zekai
Liu, Ji
Li, Guanchen
Xu, Yixing
Liu, Ziqiong
Yin, Xuanwu
Li, Dong
Barsoum, Emad
Machine Learning
Computation and Language
Structured sparsity has emerged as a popular model pruning technique, widely adopted in various architectures, including CNNs, Transformer models, and especially large language models (LLMs) in recent years. A promising direction to further improve post-pruning performance is weight permutation, which reorders model weights into patterns more amenable to pruning. However, the exponential growth of the permutation search space with the scale of Transformer architectures forces most methods to rely on greedy or heuristic algorithms, limiting the effectiveness of reordering. In this work, we propose a novel end-to-end learnable permutation framework. Our method introduces a learnable permutation cost matrix to quantify the cost of swapping any two input channels of a given weight matrix, a differentiable bipartite matching solver to obtain the optimal binary permutation matrix given a cost matrix, and a sparsity optimization loss function to directly optimize the permutation operator. We extensively validate our approach on vision and language Transformers, demonstrating that our method achieves state-of-the-art permutation results for structured sparsity.
title Learnable Permutation for Structured Sparsity on Transformer Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2601.22980