Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meyer, Marius, Kenter, Tobias, Petrica, Lucian, O'Brien, Kenneth, Blott, Michaela, Plessl, Christian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909162940661760
author Meyer, Marius
Kenter, Tobias
Petrica, Lucian
O'Brien, Kenneth
Blott, Michaela
Plessl, Christian
author_facet Meyer, Marius
Kenter, Tobias
Petrica, Lucian
O'Brien, Kenneth
Blott, Michaela
Plessl, Christian
contents Most FPGA boards in the HPC domain are well-suited for parallel scaling because of the direct integration of versatile and high-throughput network ports. However, the utilization of their network capabilities is often challenging and error-prone because the whole network stack and communication patterns have to be implemented and managed on the FPGAs. Also, this approach conceptually involves a trade-off between the performance potential of improved communication and the impact of resource consumption for communication infrastructure, since the utilized resources on the FPGAs could otherwise be used for computations. In this work, we investigate this trade-off, firstly, by using synthetic benchmarks to evaluate the different configuration options of the communication framework ACCL and their impact on communication latency and throughput. Finally, we use our findings to implement a shallow water simulation whose scalability heavily depends on low-latency communication. With a suitable configuration of ACCL, good scaling behavior can be shown to all 48 FPGAs installed in the system. Overall, the results show that the availability of inter-FPGA communication frameworks as well as the configurability of framework and network stack are crucial to achieve the best application performance with low latency communication.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18374
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
Meyer, Marius
Kenter, Tobias
Petrica, Lucian
O'Brien, Kenneth
Blott, Michaela
Plessl, Christian
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Most FPGA boards in the HPC domain are well-suited for parallel scaling because of the direct integration of versatile and high-throughput network ports. However, the utilization of their network capabilities is often challenging and error-prone because the whole network stack and communication patterns have to be implemented and managed on the FPGAs. Also, this approach conceptually involves a trade-off between the performance potential of improved communication and the impact of resource consumption for communication infrastructure, since the utilized resources on the FPGAs could otherwise be used for computations. In this work, we investigate this trade-off, firstly, by using synthetic benchmarks to evaluate the different configuration options of the communication framework ACCL and their impact on communication latency and throughput. Finally, we use our findings to implement a shallow water simulation whose scalability heavily depends on low-latency communication. With a suitable configuration of ACCL, good scaling behavior can be shown to all 48 FPGAs installed in the system. Overall, the results show that the availability of inter-FPGA communication frameworks as well as the configurability of framework and network stack are crucial to achieve the best application performance with low latency communication.
title Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2403.18374