A Proximal Operator for Inducing 2:4-Sparsity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kübler, Jonas M, Wang, Yu-Xiang, Sabach, Shoham, Ansari, Navid, Kleindessner, Matthäus, Budhathoki, Kailash, Cevher, Volkan, Karypis, George
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911123759955968
author Kübler, Jonas M
Wang, Yu-Xiang
Sabach, Shoham
Ansari, Navid
Kleindessner, Matthäus
Budhathoki, Kailash
Cevher, Volkan
Karypis, George
author_facet Kübler, Jonas M
Wang, Yu-Xiang
Sabach, Shoham
Ansari, Navid
Kleindessner, Matthäus
Budhathoki, Kailash
Cevher, Volkan
Karypis, George
contents Recent hardware advancements in AI Accelerators and GPUs allow to efficiently compute sparse matrix multiplications, especially when 2 out of 4 consecutive weights are set to zero. However, this so-called 2:4 sparsity usually comes at a decreased accuracy of the model. We derive a regularizer that exploits the local correlation of features to find better sparsity masks in trained models. We minimize the regularizer jointly with a local squared loss by deriving the proximal operator for which we show that it has an efficient solution in the 2:4-sparse case. After optimizing the mask, we use maskedgradient updates to further minimize the local squared loss. We illustrate our method on toy problems and apply it to pruning entire large language models up to 70B parameters. On models up to 13B we improve over previous state of the art algorithms, whilst on 70B models we match their performance.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18015
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Proximal Operator for Inducing 2:4-Sparsity
Kübler, Jonas M
Wang, Yu-Xiang
Sabach, Shoham
Ansari, Navid
Kleindessner, Matthäus
Budhathoki, Kailash
Cevher, Volkan
Karypis, George
Machine Learning
Recent hardware advancements in AI Accelerators and GPUs allow to efficiently compute sparse matrix multiplications, especially when 2 out of 4 consecutive weights are set to zero. However, this so-called 2:4 sparsity usually comes at a decreased accuracy of the model. We derive a regularizer that exploits the local correlation of features to find better sparsity masks in trained models. We minimize the regularizer jointly with a local squared loss by deriving the proximal operator for which we show that it has an efficient solution in the 2:4-sparse case. After optimizing the mask, we use maskedgradient updates to further minimize the local squared loss. We illustrate our method on toy problems and apply it to pruning entire large language models up to 70B parameters. On models up to 13B we improve over previous state of the art algorithms, whilst on 70B models we match their performance.
title A Proximal Operator for Inducing 2:4-Sparsity
topic Machine Learning
url https://arxiv.org/abs/2501.18015