Pack my weights and run! Minimizing overheads for in-memory computing accelerators

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Houshmand, Pouya, Verhelst, Marian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913505406353408
author Houshmand, Pouya
Verhelst, Marian
author_facet Houshmand, Pouya
Verhelst, Marian
contents In-memory computing hardware accelerators allow more than 10x improvements in peak efficiency and performance for matrix-vector multiplications (MVM) compared to conventional digital designs. For this, they have gained great interest for the acceleration of neural network workloads. Nevertheless, these potential gains are only achieved when the utilization of the computational resources is maximized and the overhead from loading operands in the memory array minimized. To this aim, this paper proposes a novel mapping algorithm for the weights in the IMC macro, based on efficient packing of the weights of network layers in the available memory. The algorithm realizes 1) minimization of weight loading times while at the same time 2) maximally exploiting the parallelism of the IMC computational fabric. A set of case studies are carried out to show achievable trade-offs for the MLPerf Tiny benchmark \cite{mlperftiny} on IMC architectures, with potential $10-100\times$ EDP improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11437
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pack my weights and run! Minimizing overheads for in-memory computing accelerators
Houshmand, Pouya
Verhelst, Marian
Hardware Architecture
Image and Video Processing
In-memory computing hardware accelerators allow more than 10x improvements in peak efficiency and performance for matrix-vector multiplications (MVM) compared to conventional digital designs. For this, they have gained great interest for the acceleration of neural network workloads. Nevertheless, these potential gains are only achieved when the utilization of the computational resources is maximized and the overhead from loading operands in the memory array minimized. To this aim, this paper proposes a novel mapping algorithm for the weights in the IMC macro, based on efficient packing of the weights of network layers in the available memory. The algorithm realizes 1) minimization of weight loading times while at the same time 2) maximally exploiting the parallelism of the IMC computational fabric. A set of case studies are carried out to show achievable trade-offs for the MLPerf Tiny benchmark \cite{mlperftiny} on IMC architectures, with potential $10-100\times$ EDP improvements.
title Pack my weights and run! Minimizing overheads for in-memory computing accelerators
topic Hardware Architecture
Image and Video Processing
url https://arxiv.org/abs/2409.11437