DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shan, Baodi, Araya-Polo, Mauricio, Chapman, Barbara
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908663003742208
author Shan, Baodi
Araya-Polo, Mauricio
Chapman, Barbara
author_facet Shan, Baodi
Araya-Polo, Mauricio
Chapman, Barbara
contents As core counts and heterogeneity rise in HPC, traditional hybrid programming models face challenges in managing distributed GPU memory and ensuring portability. This paper presents DiOMP, a distributed OpenMP framework that unifies OpenMP target offloading with the Partitioned Global Address Space (PGAS) model. Built atop LLVM/OpenMP and using GASNet-EX or GPI-2 for communication, DiOMP transparently handles global memory, supporting both symmetric and asymmetric GPU allocations. It leverages OMPCCL, a portable collective communication layer compatible with vendor libraries. DiOMP simplifies programming by abstracting device memory and communication, achieving superior scalability and programmability over traditional approaches. Evaluations on NVIDIA A100, Grace Hopper, and AMD MI250X show improved performance in micro-benchmarks and applications like matrix multiplication and Minimod, highlighting DiOMP's potential for scalable, portable, and efficient heterogeneous computing.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02486
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
Shan, Baodi
Araya-Polo, Mauricio
Chapman, Barbara
Distributed, Parallel, and Cluster Computing
As core counts and heterogeneity rise in HPC, traditional hybrid programming models face challenges in managing distributed GPU memory and ensuring portability. This paper presents DiOMP, a distributed OpenMP framework that unifies OpenMP target offloading with the Partitioned Global Address Space (PGAS) model. Built atop LLVM/OpenMP and using GASNet-EX or GPI-2 for communication, DiOMP transparently handles global memory, supporting both symmetric and asymmetric GPU allocations. It leverages OMPCCL, a portable collective communication layer compatible with vendor libraries. DiOMP simplifies programming by abstracting device memory and communication, achieving superior scalability and programmability over traditional approaches. Evaluations on NVIDIA A100, Grace Hopper, and AMD MI250X show improved performance in micro-benchmarks and applications like matrix multiplication and Minimod, highlighting DiOMP's potential for scalable, portable, and efficient heterogeneous computing.
title DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.02486