Saved in:
Bibliographic Details
Main Authors: Boža, Vladimír, Macko, Vladimír
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.18850
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908392041218048
author Boža, Vladimír
Macko, Vladimír
author_facet Boža, Vladimír
Macko, Vladimír
contents Neural networks are often challenging to work with due to their large size and complexity. To address this, various methods aim to reduce model size by sparsifying or decomposing weight matrices, such as magnitude pruning and low-rank or block-diagonal factorization. In this work, we present Double Sparse Factorization (DSF), where we factorize each weight matrix into two sparse matrices. Although solving this problem exactly is computationally infeasible, we propose an efficient heuristic based on alternating minimization via ADMM that achieves state-of-the-art results, enabling unprecedented sparsification of neural networks. For instance, in a one-shot pruning setting, our method can reduce the size of the LLaMA2-13B model by 50% while maintaining better performance than the dense LLaMA2-7B model. We also compare favorably with Optimal Brain Compression, the state-of-the-art layer-wise pruning approach for convolutional neural networks. Furthermore, accuracy improvements of our method persist even after further model fine-tuning. Code available at: https://github.com/usamec/double_sparse.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18850
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
Boža, Vladimír
Macko, Vladimír
Machine Learning
Neural networks are often challenging to work with due to their large size and complexity. To address this, various methods aim to reduce model size by sparsifying or decomposing weight matrices, such as magnitude pruning and low-rank or block-diagonal factorization. In this work, we present Double Sparse Factorization (DSF), where we factorize each weight matrix into two sparse matrices. Although solving this problem exactly is computationally infeasible, we propose an efficient heuristic based on alternating minimization via ADMM that achieves state-of-the-art results, enabling unprecedented sparsification of neural networks. For instance, in a one-shot pruning setting, our method can reduce the size of the LLaMA2-13B model by 50% while maintaining better performance than the dense LLaMA2-7B model. We also compare favorably with Optimal Brain Compression, the state-of-the-art layer-wise pruning approach for convolutional neural networks. Furthermore, accuracy improvements of our method persist even after further model fine-tuning. Code available at: https://github.com/usamec/double_sparse.
title Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
topic Machine Learning
url https://arxiv.org/abs/2409.18850