Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wahlgren, Jacob, Schieffer, Gabin, Shi, Ruimin, León, Edgar A., Pearce, Roger, Gokhale, Maya, Peng, Ivy
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909988182556672
author Wahlgren, Jacob
Schieffer, Gabin
Shi, Ruimin
León, Edgar A.
Pearce, Roger
Gokhale, Maya
Peng, Ivy
author_facet Wahlgren, Jacob
Schieffer, Gabin
Shi, Ruimin
León, Edgar A.
Pearce, Roger
Gokhale, Maya
Peng, Ivy
contents Discrete GPUs are a cornerstone of HPC and data center systems, requiring management of separate CPU and GPU memory spaces. Unified Virtual Memory (UVM) has been proposed to ease the burden of memory management; however, at a high cost in performance. The recent introduction of AMD's MI300A Accelerated Processing Units (APUs)--as deployed in the El Capitan supercomputer--enables HPC systems featuring integrated CPU and GPU with Unified Physical Memory (UPM) for the first time. This work presents the first comprehensive characterization of the UPM architecture on MI300A. We first analyze the UPM system properties, including memory latency, bandwidth, and coherence overhead. We then assess the efficiency of the system software in memory allocation, page fault handling, TLB management, and Infinity Cache utilization. We propose a set of porting strategies for transforming applications for the UPM architecture and evaluate six applications on the MI300A APU. Our results show that applications on UPM using the unified memory model can match or outperform those in the explicitly managed model--while reducing memory costs by up to 44%.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
Wahlgren, Jacob
Schieffer, Gabin
Shi, Ruimin
León, Edgar A.
Pearce, Roger
Gokhale, Maya
Peng, Ivy
Distributed, Parallel, and Cluster Computing
Performance
Discrete GPUs are a cornerstone of HPC and data center systems, requiring management of separate CPU and GPU memory spaces. Unified Virtual Memory (UVM) has been proposed to ease the burden of memory management; however, at a high cost in performance. The recent introduction of AMD's MI300A Accelerated Processing Units (APUs)--as deployed in the El Capitan supercomputer--enables HPC systems featuring integrated CPU and GPU with Unified Physical Memory (UPM) for the first time. This work presents the first comprehensive characterization of the UPM architecture on MI300A. We first analyze the UPM system properties, including memory latency, bandwidth, and coherence overhead. We then assess the efficiency of the system software in memory allocation, page fault handling, TLB management, and Infinity Cache utilization. We propose a set of porting strategies for transforming applications for the UPM architecture and evaluate six applications on the MI300A APU. Our results show that applications on UPM using the unified memory model can match or outperform those in the explicitly managed model--while reducing memory costs by up to 44%.
title Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2508.12743