Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Diehl, Patrick, Syskakis, Panagiotis, Daiß, Gregor, Brandt, Steven R., Kheirkhahan, Alireza, Singanaboina, Srinivas Yadav, Marcello, Dominic, Taylor, Chris, Leidel, John, Kaiser, Hartmut
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915102619336704
author Diehl, Patrick
Syskakis, Panagiotis
Daiß, Gregor
Brandt, Steven R.
Kheirkhahan, Alireza
Singanaboina, Srinivas Yadav
Marcello, Dominic
Taylor, Chris
Leidel, John
Kaiser, Hartmut
author_facet Diehl, Patrick
Syskakis, Panagiotis
Daiß, Gregor
Brandt, Steven R.
Kheirkhahan, Alireza
Singanaboina, Srinivas Yadav
Marcello, Dominic
Taylor, Chris
Leidel, John
Kaiser, Hartmut
contents In recent years, interest in RISC-V computing architectures has moved from academic to mainstream, especially in the field of High Performance Computing where energy limitations are increasingly a concern. As of this year, the first single board RISC-V CPUs implementing the finalized ratified vector specification are being released. The RISC-V vector specification follows in the tradition of vector processors found in the CDC STAR-100, the Cray-1, the Convex C-Series, and the NEC SX machines and accelerators. The family of vector processors offers support for variable-length array processing as opposed to the fixed-length processing functionality offered by SIMD. Vector processors offer opportunities to perform vector-chaining which allows temporary results to be used without the need to resolve memory references. In this work, we use the Octo-Tiger multi-physics, multi-scale, 3D adaptive mesh refinement astrophysics application to study these early RISC-V chips with vector machine support. We report on our experience in porting this modern C++ code (which is built upon several open-source libraries such as HPX and Kokkos) to RISC-V. In addition, we show the impact of the RISC-V Vector extension on a RISC-V single board computer by implementing the std::experimental:simd interface and integrating it with our code. We also compare the application's performance, scalability, and power consumption on desktop-grade RISC-V computer to an A64FX system.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00026
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
Diehl, Patrick
Syskakis, Panagiotis
Daiß, Gregor
Brandt, Steven R.
Kheirkhahan, Alireza
Singanaboina, Srinivas Yadav
Marcello, Dominic
Taylor, Chris
Leidel, John
Kaiser, Hartmut
Distributed, Parallel, and Cluster Computing
In recent years, interest in RISC-V computing architectures has moved from academic to mainstream, especially in the field of High Performance Computing where energy limitations are increasingly a concern. As of this year, the first single board RISC-V CPUs implementing the finalized ratified vector specification are being released. The RISC-V vector specification follows in the tradition of vector processors found in the CDC STAR-100, the Cray-1, the Convex C-Series, and the NEC SX machines and accelerators. The family of vector processors offers support for variable-length array processing as opposed to the fixed-length processing functionality offered by SIMD. Vector processors offer opportunities to perform vector-chaining which allows temporary results to be used without the need to resolve memory references. In this work, we use the Octo-Tiger multi-physics, multi-scale, 3D adaptive mesh refinement astrophysics application to study these early RISC-V chips with vector machine support. We report on our experience in porting this modern C++ code (which is built upon several open-source libraries such as HPX and Kokkos) to RISC-V. In addition, we show the impact of the RISC-V Vector extension on a RISC-V single board computer by implementing the std::experimental:simd interface and integrating it with our code. We also compare the application's performance, scalability, and power consumption on desktop-grade RISC-V computer to an A64FX system.
title Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2407.00026