NVR: Vector Runahead on NPUs for Sparse Memory Access

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hui, Zhao, Zhengpeng, Wang, Jing, Du, Yushu, Cheng, Yuan, Guo, Bing, Xiao, He, Ma, Chenhao, Han, Xiaomeng, You, Dean, Guan, Jiapeng, Wei, Ran, Yang, Dawei, Jiang, Zhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912279288610816
author Wang, Hui
Zhao, Zhengpeng
Wang, Jing
Du, Yushu
Cheng, Yuan
Guo, Bing
Xiao, He
Ma, Chenhao
Han, Xiaomeng
You, Dean
Guan, Jiapeng
Wei, Ran
Yang, Dawei
Jiang, Zhe
author_facet Wang, Hui
Zhao, Zhengpeng
Wang, Jing
Du, Yushu
Cheng, Yuan
Guo, Bing
Xiao, He
Ma, Chenhao
Han, Xiaomeng
You, Dean
Guan, Jiapeng
Wei, Ran
Yang, Dawei
Jiang, Zhe
contents Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the L2 cache size by the same amount.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13873
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NVR: Vector Runahead on NPUs for Sparse Memory Access
Wang, Hui
Zhao, Zhengpeng
Wang, Jing
Du, Yushu
Cheng, Yuan
Guo, Bing
Xiao, He
Ma, Chenhao
Han, Xiaomeng
You, Dean
Guan, Jiapeng
Wei, Ran
Yang, Dawei
Jiang, Zhe
Hardware Architecture
Artificial Intelligence
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the L2 cache size by the same amount.
title NVR: Vector Runahead on NPUs for Sparse Memory Access
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2502.13873