Addressing memory bandwidth scalability in vector processors for streaming applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Altayo, Jordi, Delestrac, Paul, Novo, David, Yang, Simey, Bhattacharjee, Debjyoti, Catthoor, Francky
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915292505964544
author Altayo, Jordi
Delestrac, Paul
Novo, David
Yang, Simey
Bhattacharjee, Debjyoti
Catthoor, Francky
author_facet Altayo, Jordi
Delestrac, Paul
Novo, David
Yang, Simey
Bhattacharjee, Debjyoti
Catthoor, Francky
contents As the size of artificial intelligence and machine learning (AI/ML) models and datasets grows, the memory bandwidth becomes a critical bottleneck. The paper presents a novel extended memory hierarchy that addresses some major memory bandwidth challenges in data-parallel AI/ML applications. While data-parallel architectures like GPUs and neural network accelerators have improved power performance compared to traditional CPUs, they can still be significantly bottlenecked by their memory bandwidth, especially when the data reuse in the loop kernels is limited. Systolic arrays (SAs) and GPUs attempt to mitigate the memory bandwidth bottleneck but can still become memory bandwidth throttled when the amount of data reuse is not sufficient to confine data access mostly to the local memories near to the processing. To mitigate this, the proposed architecture introduces three levels of on-chip memory -- local, intermediate, and global -- with an ultra-wide register and data-shufflers to improve versatility and adaptivity to varying data-parallel applications. The paper explains the innovations at a conceptual level and presents a detailed description of the architecture innovations. We also map a representative data-parallel application, like a convolutional neural network (CNN), to the proposed architecture and quantify the benefits vis-a-vis GPUs and repersentative accelerators based on systolic arrays and vector processors.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12856
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Addressing memory bandwidth scalability in vector processors for streaming applications
Altayo, Jordi
Delestrac, Paul
Novo, David
Yang, Simey
Bhattacharjee, Debjyoti
Catthoor, Francky
Hardware Architecture
As the size of artificial intelligence and machine learning (AI/ML) models and datasets grows, the memory bandwidth becomes a critical bottleneck. The paper presents a novel extended memory hierarchy that addresses some major memory bandwidth challenges in data-parallel AI/ML applications. While data-parallel architectures like GPUs and neural network accelerators have improved power performance compared to traditional CPUs, they can still be significantly bottlenecked by their memory bandwidth, especially when the data reuse in the loop kernels is limited. Systolic arrays (SAs) and GPUs attempt to mitigate the memory bandwidth bottleneck but can still become memory bandwidth throttled when the amount of data reuse is not sufficient to confine data access mostly to the local memories near to the processing. To mitigate this, the proposed architecture introduces three levels of on-chip memory -- local, intermediate, and global -- with an ultra-wide register and data-shufflers to improve versatility and adaptivity to varying data-parallel applications. The paper explains the innovations at a conceptual level and presents a detailed description of the architecture innovations. We also map a representative data-parallel application, like a convolutional neural network (CNN), to the proposed architecture and quantify the benefits vis-a-vis GPUs and repersentative accelerators based on systolic arrays and vector processors.
title Addressing memory bandwidth scalability in vector processors for streaming applications
topic Hardware Architecture
url https://arxiv.org/abs/2505.12856