Easy Acceleration with Distributed Arrays

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kepner, Jeremy, Byun, Chansup, Anderson, LaToya, Arcand, William, Bestor, David, Bergeron, William, Bonn, Alex, Burrill, Daniel, Gadepally, Vijay, Haney, Ryan, Houle, Michael, Hubbell, Matthew, Jananthan, Hayden, Jones, Michael, Luszczek, Piotr, Milechin, Lauren, Morales, Guillermo, Mullen, Julie, Prout, Andrew, Reuther, Albert, Rosa, Antonio, Yee, Charles, Michaleas, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909853831659520
author Kepner, Jeremy
Byun, Chansup
Anderson, LaToya
Arcand, William
Bestor, David
Bergeron, William
Bonn, Alex
Burrill, Daniel
Gadepally, Vijay
Haney, Ryan
Houle, Michael
Hubbell, Matthew
Jananthan, Hayden
Jones, Michael
Luszczek, Piotr
Milechin, Lauren
Morales, Guillermo
Mullen, Julie
Prout, Andrew
Reuther, Albert
Rosa, Antonio
Yee, Charles
Michaleas, Peter
author_facet Kepner, Jeremy
Byun, Chansup
Anderson, LaToya
Arcand, William
Bestor, David
Bergeron, William
Bonn, Alex
Burrill, Daniel
Gadepally, Vijay
Haney, Ryan
Houle, Michael
Hubbell, Matthew
Jananthan, Hayden
Jones, Michael
Luszczek, Piotr
Milechin, Lauren
Morales, Guillermo
Mullen, Julie
Prout, Andrew
Reuther, Albert
Rosa, Antonio
Yee, Charles
Michaleas, Peter
contents High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance while retaining productivity requires effective abstractions. Distributed arrays are one such abstraction that enables high level programming to achieve highly scalable performance. Distributed arrays achieve this performance by deriving parallelism from data locality, which naturally leads to high memory bandwidth efficiency. This paper explores distributed array performance using the STREAM memory bandwidth benchmark on a variety of hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. The hardware used spans decades and allows a direct comparison of hardware improvements for memory bandwidth over this time range; showing a 10x increase in CPU core bandwidth over 20 years, 100x increase in CPU node bandwidth over 20 years, and 5x increase in GPU node bandwidth over 5 years. Running on hundreds of MIT SuperCloud nodes simultaneously achieved a sustained bandwidth $>$1 PB/s.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Easy Acceleration with Distributed Arrays
Kepner, Jeremy
Byun, Chansup
Anderson, LaToya
Arcand, William
Bestor, David
Bergeron, William
Bonn, Alex
Burrill, Daniel
Gadepally, Vijay
Haney, Ryan
Houle, Michael
Hubbell, Matthew
Jananthan, Hayden
Jones, Michael
Luszczek, Piotr
Milechin, Lauren
Morales, Guillermo
Mullen, Julie
Prout, Andrew
Reuther, Albert
Rosa, Antonio
Yee, Charles
Michaleas, Peter
Distributed, Parallel, and Cluster Computing
Computational Engineering, Finance, and Science
Mathematical Software
Performance
High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance while retaining productivity requires effective abstractions. Distributed arrays are one such abstraction that enables high level programming to achieve highly scalable performance. Distributed arrays achieve this performance by deriving parallelism from data locality, which naturally leads to high memory bandwidth efficiency. This paper explores distributed array performance using the STREAM memory bandwidth benchmark on a variety of hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. The hardware used spans decades and allows a direct comparison of hardware improvements for memory bandwidth over this time range; showing a 10x increase in CPU core bandwidth over 20 years, 100x increase in CPU node bandwidth over 20 years, and 5x increase in GPU node bandwidth over 5 years. Running on hundreds of MIT SuperCloud nodes simultaneously achieved a sustained bandwidth $>$1 PB/s.
title Easy Acceleration with Distributed Arrays
topic Distributed, Parallel, and Cluster Computing
Computational Engineering, Finance, and Science
Mathematical Software
Performance
url https://arxiv.org/abs/2508.17493