Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Williams, Jeremy J., Liu, Felix, Trilaksono, Jordy, Tskhakaya, David, Costea, Stefan, Kos, Leon, Podolnik, Ales, Hromadka, Jakub, Hegde, Pratibha, Garcia-Gasulla, Marta, Seitz, Valentin, Jenko, Frank, Laure, Erwin, Markidis, Stefano
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909592706875392
author Williams, Jeremy J.
Liu, Felix
Trilaksono, Jordy
Tskhakaya, David
Costea, Stefan
Kos, Leon
Podolnik, Ales
Hromadka, Jakub
Hegde, Pratibha
Garcia-Gasulla, Marta
Seitz, Valentin
Jenko, Frank
Laure, Erwin
Markidis, Stefano
author_facet Williams, Jeremy J.
Liu, Felix
Trilaksono, Jordy
Tskhakaya, David
Costea, Stefan
Kos, Leon
Podolnik, Ales
Hromadka, Jakub
Hegde, Pratibha
Garcia-Gasulla, Marta
Seitz, Valentin
Jenko, Frank
Laure, Erwin
Markidis, Stefano
contents As fusion energy devices advance, plasma simulations are crucial for reactor design. Our work extends BIT1 hybrid parallelization by integrating MPI with OpenMP and OpenACC, focusing on asynchronous multi-GPU programming. Results show significant performance gains: 16 MPI ranks plus OpenMP threads reduced runtime by 53% on a petascale EuroHPC supercomputer, while OpenACC multicore achieved a 58% reduction. At 64 MPI ranks, OpenACC outperformed OpenMP, improving the particle mover function by 24%. On MareNostrum 5, OpenACC async(n) delivered strong performance, but OpenMP asynchronous multi-GPU approach proved more effective at extreme scaling, maintaining efficiency up to 400 GPUs. Speedup and parallel efficiency (PE) studies revealed OpenMP asynchronous multi-GPU achieving 8.77x speedup (54.81% PE), surpassing OpenACC (8.14x speedup, 50.87% PE). While PE declined at high node counts due to communication overhead, asynchronous execution mitigated scalability bottlenecks. OpenMP nowait and depend clauses improved GPU performance via efficient data transfer and task management. Using NVIDIA Nsight tools, we confirmed BIT1 efficiency for large-scale plasma simulations. OpenMP asynchronous multi-GPU implementation delivered exceptional performance in portability, high throughput, and GPU utilization, positioning BIT1 for exascale supercomputing and advancing fusion energy research. MareNostrum 5 brings us closer to achieving exascale performance.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
Williams, Jeremy J.
Liu, Felix
Trilaksono, Jordy
Tskhakaya, David
Costea, Stefan
Kos, Leon
Podolnik, Ales
Hromadka, Jakub
Hegde, Pratibha
Garcia-Gasulla, Marta
Seitz, Valentin
Jenko, Frank
Laure, Erwin
Markidis, Stefano
Distributed, Parallel, and Cluster Computing
Performance
Computational Physics
As fusion energy devices advance, plasma simulations are crucial for reactor design. Our work extends BIT1 hybrid parallelization by integrating MPI with OpenMP and OpenACC, focusing on asynchronous multi-GPU programming. Results show significant performance gains: 16 MPI ranks plus OpenMP threads reduced runtime by 53% on a petascale EuroHPC supercomputer, while OpenACC multicore achieved a 58% reduction. At 64 MPI ranks, OpenACC outperformed OpenMP, improving the particle mover function by 24%. On MareNostrum 5, OpenACC async(n) delivered strong performance, but OpenMP asynchronous multi-GPU approach proved more effective at extreme scaling, maintaining efficiency up to 400 GPUs. Speedup and parallel efficiency (PE) studies revealed OpenMP asynchronous multi-GPU achieving 8.77x speedup (54.81% PE), surpassing OpenACC (8.14x speedup, 50.87% PE). While PE declined at high node counts due to communication overhead, asynchronous execution mitigated scalability bottlenecks. OpenMP nowait and depend clauses improved GPU performance via efficient data transfer and task management. Using NVIDIA Nsight tools, we confirmed BIT1 efficiency for large-scale plasma simulations. OpenMP asynchronous multi-GPU implementation delivered exceptional performance in portability, high throughput, and GPU utilization, positioning BIT1 for exascale supercomputing and advancing fusion energy research. MareNostrum 5 brings us closer to achieving exascale performance.
title Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
topic Distributed, Parallel, and Cluster Computing
Performance
Computational Physics
url https://arxiv.org/abs/2404.10270