Exploiting long vectors with a CFD code: a co-design show case

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Blancafort, Marc, Ferrer, Roger, Houzeaux, Guillaume, Garcia-Gasulla, Marta, Mantovani, Filippo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909375613894656
author Blancafort, Marc
Ferrer, Roger
Houzeaux, Guillaume
Garcia-Gasulla, Marta
Mantovani, Filippo
author_facet Blancafort, Marc
Ferrer, Roger
Houzeaux, Guillaume
Garcia-Gasulla, Marta
Mantovani, Filippo
contents A current trend in HPC systems is the utilization of architectures with SIMD or vector extensions to exploit data parallelism. There are several ways to take advantage of such modern vector architectures, each with a different impact on the code and its portability. For example, the use of intrinsics, guided vectorization via pragmas, or compiler autovectorization. Our objectives are to maximize vectorization efficiency and minimize code specialization. To achieve these objectives, we rely on compiler autovectorization. We leverage a set of hardware and software tools that allow us to analyze in detail where autovectorization is suboptimal. Thus, we apply an iterative methodology that allows us to incrementally improve the efficient use of the underlying hardware. In this paper, we apply this methodology to a CFD production code. We evaluate the performance on an innovative configurable platform powered by a RISC-V core coupled with a wide vector unit capable of operating with up to 256 double precision elements. Following the vectorization process, we demonstrate a single-core speedup of 7.6$\times$ compared to its scalar implementation. Furthermore, we show that code portability is not compromised, as our solution continues to exhibit performance benefits, or at the very least, no drawbacks, on other HPC architectures such as Intel x86 and NEC SX-Aurora.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting long vectors with a CFD code: a co-design show case
Blancafort, Marc
Ferrer, Roger
Houzeaux, Guillaume
Garcia-Gasulla, Marta
Mantovani, Filippo
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
A current trend in HPC systems is the utilization of architectures with SIMD or vector extensions to exploit data parallelism. There are several ways to take advantage of such modern vector architectures, each with a different impact on the code and its portability. For example, the use of intrinsics, guided vectorization via pragmas, or compiler autovectorization. Our objectives are to maximize vectorization efficiency and minimize code specialization. To achieve these objectives, we rely on compiler autovectorization. We leverage a set of hardware and software tools that allow us to analyze in detail where autovectorization is suboptimal. Thus, we apply an iterative methodology that allows us to incrementally improve the efficient use of the underlying hardware. In this paper, we apply this methodology to a CFD production code. We evaluate the performance on an innovative configurable platform powered by a RISC-V core coupled with a wide vector unit capable of operating with up to 256 double precision elements. Following the vectorization process, we demonstrate a single-core speedup of 7.6$\times$ compared to its scalar implementation. Furthermore, we show that code portability is not compromised, as our solution continues to exhibit performance benefits, or at the very least, no drawbacks, on other HPC architectures such as Intel x86 and NEC SX-Aurora.
title Exploiting long vectors with a CFD code: a co-design show case
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
url https://arxiv.org/abs/2411.00815