Easy Acceleration with Distributed Arrays
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909853831659520 |
|---|---|
| author | Kepner, Jeremy Byun, Chansup Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel Gadepally, Vijay Haney, Ryan Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Luszczek, Piotr Milechin, Lauren Morales, Guillermo Mullen, Julie Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Michaleas, Peter |
| author_facet | Kepner, Jeremy Byun, Chansup Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel Gadepally, Vijay Haney, Ryan Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Luszczek, Piotr Milechin, Lauren Morales, Guillermo Mullen, Julie Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Michaleas, Peter |
| contents | High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance while retaining productivity requires effective abstractions. Distributed arrays are one such abstraction that enables high level programming to achieve highly scalable performance. Distributed arrays achieve this performance by deriving parallelism from data locality, which naturally leads to high memory bandwidth efficiency. This paper explores distributed array performance using the STREAM memory bandwidth benchmark on a variety of hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. The hardware used spans decades and allows a direct comparison of hardware improvements for memory bandwidth over this time range; showing a 10x increase in CPU core bandwidth over 20 years, 100x increase in CPU node bandwidth over 20 years, and 5x increase in GPU node bandwidth over 5 years. Running on hundreds of MIT SuperCloud nodes simultaneously achieved a sustained bandwidth $>$1 PB/s. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_17493 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Easy Acceleration with Distributed Arrays Kepner, Jeremy Byun, Chansup Anderson, LaToya Arcand, William Bestor, David Bergeron, William Bonn, Alex Burrill, Daniel Gadepally, Vijay Haney, Ryan Houle, Michael Hubbell, Matthew Jananthan, Hayden Jones, Michael Luszczek, Piotr Milechin, Lauren Morales, Guillermo Mullen, Julie Prout, Andrew Reuther, Albert Rosa, Antonio Yee, Charles Michaleas, Peter Distributed, Parallel, and Cluster Computing Computational Engineering, Finance, and Science Mathematical Software Performance High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance while retaining productivity requires effective abstractions. Distributed arrays are one such abstraction that enables high level programming to achieve highly scalable performance. Distributed arrays achieve this performance by deriving parallelism from data locality, which naturally leads to high memory bandwidth efficiency. This paper explores distributed array performance using the STREAM memory bandwidth benchmark on a variety of hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. The hardware used spans decades and allows a direct comparison of hardware improvements for memory bandwidth over this time range; showing a 10x increase in CPU core bandwidth over 20 years, 100x increase in CPU node bandwidth over 20 years, and 5x increase in GPU node bandwidth over 5 years. Running on hundreds of MIT SuperCloud nodes simultaneously achieved a sustained bandwidth $>$1 PB/s. |
| title | Easy Acceleration with Distributed Arrays |
| topic | Distributed, Parallel, and Cluster Computing Computational Engineering, Finance, and Science Mathematical Software Performance |
| url | https://arxiv.org/abs/2508.17493 |