Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fusco, Luigi, Khalilov, Mikhail, Chrapek, Marcin, Chukkapalli, Giridhar, Schulthess, Thomas, Hoefler, Torsten
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913480144060416
author Fusco, Luigi
Khalilov, Mikhail
Chrapek, Marcin
Chukkapalli, Giridhar
Schulthess, Thomas
Hoefler, Torsten
author_facet Fusco, Luigi
Khalilov, Mikhail
Chrapek, Marcin
Chukkapalli, Giridhar
Schulthess, Thomas
Hoefler, Torsten
contents Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and climate modeling. With many workloads becoming memory-bound, improving the communication latency and bandwidth within the system has become a main driver in the development of new architectures. The Grace Hopper Superchip (GH200) is a significant step in the direction of tightly coupled heterogeneous systems, in which all CPUs and GPUs share a unified address space and support transparent fine grained access to all main memory on the system. We characterize both intra- and inter-node memory operations on the Quad GH200 nodes of the new Swiss National Supercomputing Centre Alps supercomputer, and show the importance of careful memory placement on example workloads, highlighting tradeoffs and opportunities.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11556
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Fusco, Luigi
Khalilov, Mikhail
Chrapek, Marcin
Chukkapalli, Giridhar
Schulthess, Thomas
Hoefler, Torsten
Distributed, Parallel, and Cluster Computing
Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and climate modeling. With many workloads becoming memory-bound, improving the communication latency and bandwidth within the system has become a main driver in the development of new architectures. The Grace Hopper Superchip (GH200) is a significant step in the direction of tightly coupled heterogeneous systems, in which all CPUs and GPUs share a unified address space and support transparent fine grained access to all main memory on the system. We characterize both intra- and inter-node memory operations on the Quad GH200 nodes of the new Swiss National Supercomputing Centre Alps supercomputer, and show the importance of careful memory placement on example workloads, highlighting tradeoffs and opportunities.
title Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.11556