Do deep neural networks utilize the weight space efficiently?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koyun, Onur Can, Töreyin, Behçet Uğur
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913214862721024
author Koyun, Onur Can
Töreyin, Behçet Uğur
author_facet Koyun, Onur Can
Töreyin, Behçet Uğur
contents Deep learning models like Transformers and Convolutional Neural Networks (CNNs) have revolutionized various domains, but their parameter-intensive nature hampers deployment in resource-constrained settings. In this paper, we introduce a novel concept utilizes column space and row space of weight matrices, which allows for a substantial reduction in model parameters without compromising performance. Leveraging this paradigm, we achieve parameter-efficient deep learning models.. Our approach applies to both Bottleneck and Attention layers, effectively halving the parameters while incurring only minor performance degradation. Extensive experiments conducted on the ImageNet dataset with ViT and ResNet50 demonstrate the effectiveness of our method, showcasing competitive performance when compared to traditional models. This approach not only addresses the pressing demand for parameter efficient deep learning solutions but also holds great promise for practical deployment in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16438
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do deep neural networks utilize the weight space efficiently?
Koyun, Onur Can
Töreyin, Behçet Uğur
Machine Learning
Artificial Intelligence
Deep learning models like Transformers and Convolutional Neural Networks (CNNs) have revolutionized various domains, but their parameter-intensive nature hampers deployment in resource-constrained settings. In this paper, we introduce a novel concept utilizes column space and row space of weight matrices, which allows for a substantial reduction in model parameters without compromising performance. Leveraging this paradigm, we achieve parameter-efficient deep learning models.. Our approach applies to both Bottleneck and Attention layers, effectively halving the parameters while incurring only minor performance degradation. Extensive experiments conducted on the ImageNet dataset with ViT and ResNet50 demonstrate the effectiveness of our method, showcasing competitive performance when compared to traditional models. This approach not only addresses the pressing demand for parameter efficient deep learning solutions but also holds great promise for practical deployment in real-world scenarios.
title Do deep neural networks utilize the weight space efficiently?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2401.16438