FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shu, Zhihao, Sanim, Md Musfiqur Rahman, Zheng, Hangyu, Zhu, Kunxiong, Yin, Miao, Agrawal, Gagan, Niu, Wei
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914334196629504
author Shu, Zhihao
Sanim, Md Musfiqur Rahman
Zheng, Hangyu
Zhu, Kunxiong
Yin, Miao
Agrawal, Gagan
Niu, Wei
author_facet Shu, Zhihao
Sanim, Md Musfiqur Rahman
Zheng, Hangyu
Zhu, Kunxiong
Yin, Miao
Agrawal, Gagan
Niu, Wei
contents The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0x to 8.4x memory reduction and 1.7x to 75.0x speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
Shu, Zhihao
Sanim, Md Musfiqur Rahman
Zheng, Hangyu
Zhu, Kunxiong
Yin, Miao
Agrawal, Gagan
Niu, Wei
Distributed, Parallel, and Cluster Computing
Machine Learning
The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0x to 8.4x memory reduction and 1.7x to 75.0x speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs.
title FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2602.15379