Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Han, Yunchu, Nan, Zhaojun, Zhou, Sheng, Niu, Zhisheng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918149637537792
author Han, Yunchu
Nan, Zhaojun
Zhou, Sheng
Niu, Zhisheng
author_facet Han, Yunchu
Nan, Zhaojun
Zhou, Sheng
Niu, Zhisheng
contents Deep neural networks (DNNs) have been widely applied in diverse applications, but the problems of high latency and energy overhead are inevitable on resource-constrained devices. To address this challenge, most researchers focus on the dynamic voltage and frequency scaling (DVFS) technique to balance the latency and energy consumption by changing the computing frequency of processors. However, the adjustment of memory frequency is usually ignored and not fully utilized to achieve efficient DNN inference, which also plays a significant role in the inference time and energy consumption. In this paper, we first investigate the impact of joint memory frequency and computing frequency scaling on the inference time and energy consumption with a model-based and data-driven method. Then by combining with the fitting parameters of different DNN models, we give a preliminary analysis for the proposed model to see the effects of adjusting memory frequency and computing frequency simultaneously. Finally, simulation results in local inference and cooperative inference cases further validate the effectiveness of jointly scaling the memory frequency and computing frequency to reduce the energy consumption of devices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
Han, Yunchu
Nan, Zhaojun
Zhou, Sheng
Niu, Zhisheng
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Deep neural networks (DNNs) have been widely applied in diverse applications, but the problems of high latency and energy overhead are inevitable on resource-constrained devices. To address this challenge, most researchers focus on the dynamic voltage and frequency scaling (DVFS) technique to balance the latency and energy consumption by changing the computing frequency of processors. However, the adjustment of memory frequency is usually ignored and not fully utilized to achieve efficient DNN inference, which also plays a significant role in the inference time and energy consumption. In this paper, we first investigate the impact of joint memory frequency and computing frequency scaling on the inference time and energy consumption with a model-based and data-driven method. Then by combining with the fitting parameters of different DNN models, we give a preliminary analysis for the proposed model to see the effects of adjusting memory frequency and computing frequency simultaneously. Finally, simulation results in local inference and cooperative inference cases further validate the effectiveness of jointly scaling the memory frequency and computing frequency to reduce the energy consumption of devices.
title Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17970