Can Tensor Cores Benefit Memory-Bound Kernels? (No!)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lingqi, Huang, Jiajun, Di, Sheng, Matsuoka, Satoshi, Wahib, Mohamed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912254809604096
author Zhang, Lingqi
Huang, Jiajun
Di, Sheng
Matsuoka, Satoshi
Wahib, Mohamed
author_facet Zhang, Lingqi
Huang, Jiajun
Di, Sheng
Matsuoka, Satoshi
Wahib, Mohamed
contents Tensor cores are specialized processing units within GPUs that have demonstrated significant efficiency gains in compute-bound applications such as Deep Learning Training by accelerating dense matrix operations. Given their success, researchers have attempted to extend tensor core capabilities beyond dense matrix computations to other computational patterns, including memory-bound kernels. Recent studies have reported that tensor cores can outperform traditional CUDA cores even on memory-bound kernels, where the primary performance bottleneck is not computation. In this research, we challenge these findings through both theoretical and empirical analysis. Our theoretical analysis reveals that tensor cores can achieve a maximum speedup of only 1.33x over CUDA cores for memory-bound kernels in double precision (for V100, A100, and H100 GPUs). We validate this theoretical limit through empirical analysis of three representative memory-bound kernels-STREAM Scale, SpMV, and stencil. We demonstrate that optimizing memory-bound kernels using tensor cores does not yield sound performance improvements over CUDA cores.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
Zhang, Lingqi
Huang, Jiajun
Di, Sheng
Matsuoka, Satoshi
Wahib, Mohamed
Distributed, Parallel, and Cluster Computing
Performance
Tensor cores are specialized processing units within GPUs that have demonstrated significant efficiency gains in compute-bound applications such as Deep Learning Training by accelerating dense matrix operations. Given their success, researchers have attempted to extend tensor core capabilities beyond dense matrix computations to other computational patterns, including memory-bound kernels. Recent studies have reported that tensor cores can outperform traditional CUDA cores even on memory-bound kernels, where the primary performance bottleneck is not computation. In this research, we challenge these findings through both theoretical and empirical analysis. Our theoretical analysis reveals that tensor cores can achieve a maximum speedup of only 1.33x over CUDA cores for memory-bound kernels in double precision (for V100, A100, and H100 GPUs). We validate this theoretical limit through empirical analysis of three representative memory-bound kernels-STREAM Scale, SpMV, and stencil. We demonstrate that optimizing memory-bound kernels using tensor cores does not yield sound performance improvements over CUDA cores.
title Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2502.16851