Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Weile, Fan, Ruibo, Li, Zeyu, Du, Dayou, Liu, Hongyuan, Wang, Qiang, Chu, Xiaowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914021014241280
author Luo, Weile
Fan, Ruibo
Li, Zeyu
Du, Dayou
Liu, Hongyuan
Wang, Qiang
Chu, Xiaowen
author_facet Luo, Weile
Fan, Ruibo
Li, Zeyu
Du, Dayou
Liu, Hongyuan
Wang, Qiang
Chu, Xiaowen
contents This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper's memory subsystem, highlighting improvements in the L2 partitioned cache and global memory access compared to Ampere and Ada Lovelace. The evaluation of Hopper's fourth-generation tensor cores reveals the benefits of FP8 precision and asynchronous wgmma instructions for matrix operations. Additionally, we investigate the performance of DPX instructions for dynamic programming, distributed shared memory (DSM) for inter-SM communication, and the Tensor Memory Accelerator (TMA) for asynchronous data movement. Through multi-level evaluation, we discover that the Hopper architecture demonstrates significant acceleration potential in real-world applications. For instance, the asynchronous programming model supported by TMA achieves a 1.5x speedup in matrix multiplication, FP8 delivers nearly double the performance of FP16, and DPX instructions accelerate a computational biology algorithm by at least 4.75x. Our findings provide actionable insights for optimizing compute-intensive workloads, from AI training to bioinformatics, on Hopper GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Luo, Weile
Fan, Ruibo
Li, Zeyu
Du, Dayou
Liu, Hongyuan
Wang, Qiang
Chu, Xiaowen
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper's memory subsystem, highlighting improvements in the L2 partitioned cache and global memory access compared to Ampere and Ada Lovelace. The evaluation of Hopper's fourth-generation tensor cores reveals the benefits of FP8 precision and asynchronous wgmma instructions for matrix operations. Additionally, we investigate the performance of DPX instructions for dynamic programming, distributed shared memory (DSM) for inter-SM communication, and the Tensor Memory Accelerator (TMA) for asynchronous data movement. Through multi-level evaluation, we discover that the Hopper architecture demonstrates significant acceleration potential in real-world applications. For instance, the asynchronous programming model supported by TMA achieves a 1.5x speedup in matrix multiplication, FP8 delivers nearly double the performance of FP16, and DPX instructions accelerate a computational biology algorithm by at least 4.75x. Our findings provide actionable insights for optimizing compute-intensive workloads, from AI training to bioinformatics, on Hopper GPUs.
title Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
Performance
url https://arxiv.org/abs/2501.12084