Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Changwoo, Kim, Hun-Seok
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910358107586560
author Lee, Changwoo
Kim, Hun-Seok
author_facet Lee, Changwoo
Kim, Hun-Seok
contents This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge arises because the optimal weight matrix structure in popular neural network models is obscure in most cases and may vary from layer to layer even in the same network. Prior structured matrices proposed for efficient DNNs were mostly hand-crafted without a generalized framework to systematically learn them. To address this issue, we propose a generalized and differentiable framework to learn efficient structures of weight matrices by gradient descent. We first define a new class of structured matrices that covers a wide range of structured matrices in the literature by adjusting the structural parameters. Then, the frequency-domain differentiable parameterization scheme based on the Gaussian-Dirichlet kernel is adopted to learn the structural parameters by proximal gradient descent. On the image and language tasks, our method learns efficient DNNs with structured matrices, achieving lower complexity and/or higher performance than prior approaches that employ low-rank, block-sparse, or block-low-rank matrices.
format Preprint
id arxiv_https___arxiv_org_abs_2310_18882
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
Lee, Changwoo
Kim, Hun-Seok
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Image and Video Processing
Signal Processing
This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge arises because the optimal weight matrix structure in popular neural network models is obscure in most cases and may vary from layer to layer even in the same network. Prior structured matrices proposed for efficient DNNs were mostly hand-crafted without a generalized framework to systematically learn them. To address this issue, we propose a generalized and differentiable framework to learn efficient structures of weight matrices by gradient descent. We first define a new class of structured matrices that covers a wide range of structured matrices in the literature by adjusting the structural parameters. Then, the frequency-domain differentiable parameterization scheme based on the Gaussian-Dirichlet kernel is adopted to learn the structural parameters by proximal gradient descent. On the image and language tasks, our method learns efficient DNNs with structured matrices, achieving lower complexity and/or higher performance than prior approaches that employ low-rank, block-sparse, or block-low-rank matrices.
title Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Image and Video Processing
Signal Processing
url https://arxiv.org/abs/2310.18882