LayerCollapse: Adaptive compression of neural networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shabgahi, Soheil Zibakhsh, Shariff, Mohammad Sohail, Koushanfar, Farinaz
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909374801248256
author Shabgahi, Soheil Zibakhsh
Shariff, Mohammad Sohail
Koushanfar, Farinaz
author_facet Shabgahi, Soheil Zibakhsh
Shariff, Mohammad Sohail
Koushanfar, Farinaz
contents Handling the ever-increasing scale of contemporary deep learning and transformer-based models poses a significant challenge. Overparameterized Transformer networks outperform prior art in Natural Language processing and Computer Vision. These models contain hundreds of millions of parameters, demanding significant computational resources and making them prone to overfitting on down stream tasks. In this work we present LayerCollapse, a novel structured pruning method to reduce the depth of fully connected layers. We propose an innovative regularizer that promotes shallow fully connected layers, compressing the model with minimal performance impact. This regularizer enables post-training compression without fine-tuning while preserving performance. LayerCollapse controls model expressiveness by regularizing the activation functions between fully connected layers, modulating them to linearity. A linear activation function collapses the rank of a transformation to the rank of the corresponding linear transformation, which demands less resources from the hardware. We demonstrate the effectiveness of LayerCollapse by showing its compression capabilities in sentimental analysis, text generation, and image classification benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17943
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LayerCollapse: Adaptive compression of neural networks
Shabgahi, Soheil Zibakhsh
Shariff, Mohammad Sohail
Koushanfar, Farinaz
Machine Learning
Artificial Intelligence
Handling the ever-increasing scale of contemporary deep learning and transformer-based models poses a significant challenge. Overparameterized Transformer networks outperform prior art in Natural Language processing and Computer Vision. These models contain hundreds of millions of parameters, demanding significant computational resources and making them prone to overfitting on down stream tasks. In this work we present LayerCollapse, a novel structured pruning method to reduce the depth of fully connected layers. We propose an innovative regularizer that promotes shallow fully connected layers, compressing the model with minimal performance impact. This regularizer enables post-training compression without fine-tuning while preserving performance. LayerCollapse controls model expressiveness by regularizing the activation functions between fully connected layers, modulating them to linearity. A linear activation function collapses the rank of a transformation to the rank of the corresponding linear transformation, which demands less resources from the hardware. We demonstrate the effectiveness of LayerCollapse by showing its compression capabilities in sentimental analysis, text generation, and image classification benchmarks.
title LayerCollapse: Adaptive compression of neural networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2311.17943