Portable Lattice QCD implementation based on OpenCL

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kumar, Piyush, Borsanyi, Szabolcs, Guenther, Jana N., Wong, Chik Him
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929699089809408
author Kumar, Piyush
Borsanyi, Szabolcs
Guenther, Jana N.
Wong, Chik Him
author_facet Kumar, Piyush
Borsanyi, Szabolcs
Guenther, Jana N.
Wong, Chik Him
contents The presence of GPU from different vendors demands the Lattice QCD codes to support multiple architectures. To this end, Open Computing Language (OpenCL) is one of the viable frameworks for writing a portable code. It is of interest to find out how the OpenCL implementation performs as compared to the code based on a dedicated programming interface such as CUDA for Nvidia GPUs. We have developed an OpenCL backend for our already existing code of the Wuppertal-Budapest collaboration. In this contribution, we show benchmarks of the most time consuming part of the numerical simulation, namely, the inversion of the Dirac operator. We present the code performance on the JUWELS and LUMI Supercomputers based on Nvidia and AMD graphics cards, respectively, and compare with the CUDA backend implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03249
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Portable Lattice QCD implementation based on OpenCL
Kumar, Piyush
Borsanyi, Szabolcs
Guenther, Jana N.
Wong, Chik Him
High Energy Physics - Lattice
The presence of GPU from different vendors demands the Lattice QCD codes to support multiple architectures. To this end, Open Computing Language (OpenCL) is one of the viable frameworks for writing a portable code. It is of interest to find out how the OpenCL implementation performs as compared to the code based on a dedicated programming interface such as CUDA for Nvidia GPUs. We have developed an OpenCL backend for our already existing code of the Wuppertal-Budapest collaboration. In this contribution, we show benchmarks of the most time consuming part of the numerical simulation, namely, the inversion of the Dirac operator. We present the code performance on the JUWELS and LUMI Supercomputers based on Nvidia and AMD graphics cards, respectively, and compare with the CUDA backend implementation.
title Portable Lattice QCD implementation based on OpenCL
topic High Energy Physics - Lattice
url https://arxiv.org/abs/2502.03249