Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hajimolahoseini, Habib, Ahmed, Walid, Liu, Yang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912390280380416
author Hajimolahoseini, Habib
Ahmed, Walid
Liu, Yang
author_facet Hajimolahoseini, Habib
Ahmed, Walid
Liu, Yang
contents Low Rank Decomposition (LRD) is a model compression technique applied to the weight tensors of deep learning models in order to reduce the number of trainable parameters and computational complexity. However, due to high number of new layers added to the architecture after applying LRD, it may not lead to a high training/inference acceleration if the decomposition ranks are not small enough. The issue is that using small ranks increases the risk of significant accuracy drop after decomposition. In this paper, we propose two techniques for accelerating low rank decomposed models without requiring to use small ranks for decomposition. These methods include rank optimization and sequential freezing of decomposed layers. We perform experiments on both convolutional and transformer-based models. Experiments show that these techniques can improve the model throughput up to 60% during training and 37% during inference when combined together while preserving the accuracy close to that of the original models
format Preprint
id arxiv_https___arxiv_org_abs_2309_03824
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
Hajimolahoseini, Habib
Ahmed, Walid
Liu, Yang
Machine Learning
Artificial Intelligence
Computation and Language
Low Rank Decomposition (LRD) is a model compression technique applied to the weight tensors of deep learning models in order to reduce the number of trainable parameters and computational complexity. However, due to high number of new layers added to the architecture after applying LRD, it may not lead to a high training/inference acceleration if the decomposition ranks are not small enough. The issue is that using small ranks increases the risk of significant accuracy drop after decomposition. In this paper, we propose two techniques for accelerating low rank decomposed models without requiring to use small ranks for decomposition. These methods include rank optimization and sequential freezing of decomposed layers. We perform experiments on both convolutional and transformer-based models. Experiments show that these techniques can improve the model throughput up to 60% during training and 37% during inference when combined together while preserving the accuracy close to that of the original models
title Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2309.03824