CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910943052562432 |
|---|---|
| author | Cai, Tianhao Wang, Liang Xiao, Limin Han, Meng Wang, Zeyu Sun, Lin Liao, Xiaojian |
| author_facet | Cai, Tianhao Wang, Liang Xiao, Limin Han, Meng Wang, Zeyu Sun, Lin Liao, Xiaojian |
| contents | With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56$\times$ (1.88$\times$ on average). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_06625 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs Cai, Tianhao Wang, Liang Xiao, Limin Han, Meng Wang, Zeyu Sun, Lin Liao, Xiaojian Hardware Architecture Artificial Intelligence Operating Systems With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56$\times$ (1.88$\times$ on average). |
| title | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs |
| topic | Hardware Architecture Artificial Intelligence Operating Systems |
| url | https://arxiv.org/abs/2505.06625 |