Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Zhu, Sun, Yu, Parakal, Dhatri, Fang, Bo, Farrell, Steven, Bauer, Gregory H., Bode, Brett, Foster, Ian T., Papka, Michael E., Gropp, William, Zhang, Zhao, Yang, Lishan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2508.03513
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909770114400256
author Zhu, Zhu
Sun, Yu
Parakal, Dhatri
Fang, Bo
Farrell, Steven
Bauer, Gregory H.
Bode, Brett
Foster, Ian T.
Papka, Michael E.
Gropp, William
Zhang, Zhao
Yang, Lishan
author_facet Zhu, Zhu
Sun, Yu
Parakal, Dhatri
Fang, Bo
Farrell, Steven
Bauer, Gregory H.
Bode, Brett
Foster, Ian T.
Papka, Michael E.
Gropp, William
Zhang, Zhao
Yang, Lishan
contents Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an essential step toward achieving efficient and reliable HPC systems. In this work, we present a large-scale cross-supercomputer study to characterize GPU memory reliability, covering three supercomputers - Delta, Polaris, and Perlmutter - all equipped with NVIDIA A100 GPUs. We examine error logs spanning 67.77 million GPU device-hours across 10,693 GPUs. We compare error rates and mean-time-between-errors (MTBE) and highlight both shared and distinct error characteristics among these three systems. Based on these observations and analyses, we discuss the implications and lessons learned, focusing on the reliable operation of supercomputers, the choice of checkpointing interval, and the comparison of reliability characteristics with those of previous-generation GPUs. Our characterization study provides valuable insights into fault-tolerant HPC system design and operation, enabling more efficient execution of HPC applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the Landscape of Ampere GPU Memory Errors
Zhu, Zhu
Sun, Yu
Parakal, Dhatri
Fang, Bo
Farrell, Steven
Bauer, Gregory H.
Bode, Brett
Foster, Ian T.
Papka, Michael E.
Gropp, William
Zhang, Zhao
Yang, Lishan
Distributed, Parallel, and Cluster Computing
Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an essential step toward achieving efficient and reliable HPC systems. In this work, we present a large-scale cross-supercomputer study to characterize GPU memory reliability, covering three supercomputers - Delta, Polaris, and Perlmutter - all equipped with NVIDIA A100 GPUs. We examine error logs spanning 67.77 million GPU device-hours across 10,693 GPUs. We compare error rates and mean-time-between-errors (MTBE) and highlight both shared and distinct error characteristics among these three systems. Based on these observations and analyses, we discuss the implications and lessons learned, focusing on the reliable operation of supercomputers, the choice of checkpointing interval, and the comparison of reliability characteristics with those of previous-generation GPUs. Our characterization study provides valuable insights into fault-tolerant HPC system design and operation, enabling more efficient execution of HPC applications.
title Understanding the Landscape of Ampere GPU Memory Errors
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.03513