GPU Acceleration of Monte Carlo Tallies on Unstructured Meshes in OpenMC with PUMI-Tally

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hasan, Fuad, Smith, Cameron W., Shephard, Mark S., Churchill, R. Michael, Wilkie, George J., Romano, Paul K., Shriwise, Patrick C., Merson, Jacob S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908339912310784
author Hasan, Fuad
Smith, Cameron W.
Shephard, Mark S.
Churchill, R. Michael
Wilkie, George J.
Romano, Paul K.
Shriwise, Patrick C.
Merson, Jacob S.
author_facet Hasan, Fuad
Smith, Cameron W.
Shephard, Mark S.
Churchill, R. Michael
Wilkie, George J.
Romano, Paul K.
Shriwise, Patrick C.
Merson, Jacob S.
contents Unstructured mesh tallies are a bottleneck in Monte Carlo neutral particle transport simulations of fusion reactors. This paper introduces the PUMI-Tally library that takes advantage of mesh adjacency information to accelerate these tallies on CPUs and GPUs. For a fixed source simulation using track-length tallies, we achieved a speed-up of 19.7X on an NVIDIA A100, and 9.2X using OpenMP on 128 threads of two AMD EPYC 7763 CPUs on NERSC Perlmutter. On the Empire AI alpha system, we achieved a speed-up of 20X using an NVIDIA H100 and 96 threads of an Intel Xenon 8568Y+. Our method showed better scaling with number of particles and number of elements. Additionally, we observed a 199X reduction in the number of allocations during initialization and the first three iterations, with a similar overall memory consumption. And, our hybrid CPU/GPU method demonstrated a 6.69X improvement in the energy consumption over the current approach.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPU Acceleration of Monte Carlo Tallies on Unstructured Meshes in OpenMC with PUMI-Tally
Hasan, Fuad
Smith, Cameron W.
Shephard, Mark S.
Churchill, R. Michael
Wilkie, George J.
Romano, Paul K.
Shriwise, Patrick C.
Merson, Jacob S.
Distributed, Parallel, and Cluster Computing
Computational Physics
Unstructured mesh tallies are a bottleneck in Monte Carlo neutral particle transport simulations of fusion reactors. This paper introduces the PUMI-Tally library that takes advantage of mesh adjacency information to accelerate these tallies on CPUs and GPUs. For a fixed source simulation using track-length tallies, we achieved a speed-up of 19.7X on an NVIDIA A100, and 9.2X using OpenMP on 128 threads of two AMD EPYC 7763 CPUs on NERSC Perlmutter. On the Empire AI alpha system, we achieved a speed-up of 20X using an NVIDIA H100 and 96 threads of an Intel Xenon 8568Y+. Our method showed better scaling with number of particles and number of elements. Additionally, we observed a 199X reduction in the number of allocations during initialization and the first three iterations, with a similar overall memory consumption. And, our hybrid CPU/GPU method demonstrated a 6.69X improvement in the energy consumption over the current approach.
title GPU Acceleration of Monte Carlo Tallies on Unstructured Meshes in OpenMC with PUMI-Tally
topic Distributed, Parallel, and Cluster Computing
Computational Physics
url https://arxiv.org/abs/2504.19048