CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Tianhao, Wang, Liang, Xiao, Limin, Han, Meng, Wang, Zeyu, Sun, Lin, Liao, Xiaojian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910943052562432
author Cai, Tianhao
Wang, Liang
Xiao, Limin
Han, Meng
Wang, Zeyu
Sun, Lin
Liao, Xiaojian
author_facet Cai, Tianhao
Wang, Liang
Xiao, Limin
Han, Meng
Wang, Zeyu
Sun, Lin
Liao, Xiaojian
contents With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56$\times$ (1.88$\times$ on average).
format Preprint
id arxiv_https___arxiv_org_abs_2505_06625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
Cai, Tianhao
Wang, Liang
Xiao, Limin
Han, Meng
Wang, Zeyu
Sun, Lin
Liao, Xiaojian
Hardware Architecture
Artificial Intelligence
Operating Systems
With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56$\times$ (1.88$\times$ on average).
title CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
topic Hardware Architecture
Artificial Intelligence
Operating Systems
url https://arxiv.org/abs/2505.06625