DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Yunchu, Nan, Zhaojun, Zhou, Sheng, Niu, Zhisheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
by: Nan, Zhaojun, et al.
Published: (2025)
by: Nan, Zhaojun, et al.
Published: (2025)
Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
by: Han, Yunchu, et al.
Published: (2025)
by: Han, Yunchu, et al.
Published: (2025)
C-MASS: Combinatorial Mobility-Aware Sensor Scheduling for Collaborative Perception with Second-Order Topology Approximation
by: Jia, Yukuan, et al.
Published: (2024)
by: Jia, Yukuan, et al.
Published: (2024)
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
by: Cheng, Zhipeng, et al.
Published: (2025)
by: Cheng, Zhipeng, et al.
Published: (2025)
Velocity and Density-Aware RRI Analysis and Optimization for AoI Minimization in IoV SPS
by: Ji, Maoxin, et al.
Published: (2025)
by: Ji, Maoxin, et al.
Published: (2025)
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
by: Liu, Renyuan, et al.
Published: (2025)
by: Liu, Renyuan, et al.
Published: (2025)
DNN-Enabled Multi-User Beamforming for Throughput Maximization under Adjustable Fairness
by: Lu, Kaifeng, et al.
Published: (2026)
by: Lu, Kaifeng, et al.
Published: (2026)
Adaptive Compression-Aware Split Learning and Inference for Enhanced Network Efficiency
by: Mudvari, Akrit, et al.
Published: (2023)
by: Mudvari, Akrit, et al.
Published: (2023)
SwiftQueue: Optimizing Low-Latency Applications with Swift Packet Queuing
by: Ray, Siddhant, et al.
Published: (2024)
by: Ray, Siddhant, et al.
Published: (2024)
A Constrained RL Approach for Cost-Efficient Delivery of Latency-Sensitive Applications
by: Aygün, Ozan, et al.
Published: (2026)
by: Aygün, Ozan, et al.
Published: (2026)
Jamming-Resilient PRB Reservation for Latency-Critical O-RAN Network Slicing
by: Delavari, Elahe, et al.
Published: (2026)
by: Delavari, Elahe, et al.
Published: (2026)
Toward Enhanced Reinforcement Learning-Based Resource Management via Digital Twin: Opportunities, Applications, and Challenges
by: Cheng, Nan, et al.
Published: (2024)
by: Cheng, Nan, et al.
Published: (2024)
Context-Aware Mobile Network Performance Prediction Using Network & Remote Sensing Data
by: Shibli, Ali, et al.
Published: (2024)
by: Shibli, Ali, et al.
Published: (2024)
Multi-Plane HyperX: A Low-Latency and Cost-Effective Network for Large-Scale AI and HPC Systems
by: Wang, Ziyu, et al.
Published: (2026)
by: Wang, Ziyu, et al.
Published: (2026)
Inference-to-complete: A High-performance and Programmable Data-plane Co-processor for Neural-network-driven Traffic Analysis
by: Wen, Dong, et al.
Published: (2024)
by: Wen, Dong, et al.
Published: (2024)
Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling
by: Li, Jialong, et al.
Published: (2024)
by: Li, Jialong, et al.
Published: (2024)
PACC: Protocol-Aware Cross-Layer Compression for Compact Network Traffic Representation
by: Guo, Zhaochen, et al.
Published: (2026)
by: Guo, Zhaochen, et al.
Published: (2026)
MamNet: A Novel Hybrid Model for Time-Series Forecasting and Frequency Pattern Analysis in Network Traffic
by: Zhang, Yujun, et al.
Published: (2025)
by: Zhang, Yujun, et al.
Published: (2025)
Prioritizing Latency with Profit: A DRL-Based Admission Control for 5G Network Slices
by: Chakraborty, Proggya, et al.
Published: (2025)
by: Chakraborty, Proggya, et al.
Published: (2025)
NeuroRisk: Physics-Informed Neural Optimization for Risk-Aware Traffic Engineering
by: Mao, Yingming, et al.
Published: (2026)
by: Mao, Yingming, et al.
Published: (2026)
Distributed Learning and Inference Systems: A Networking Perspective
by: Moussa, Hesham G., et al.
Published: (2025)
by: Moussa, Hesham G., et al.
Published: (2025)
Diffusion Models as Network Optimizers: Explorations and Analysis
by: Liang, Ruihuai, et al.
Published: (2024)
by: Liang, Ruihuai, et al.
Published: (2024)
Cloud-Edge Collaborative Large Models for Robust Photovoltaic Power Forecasting
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Latency-Aware Inter-domain Routing
by: Lin, Shihan, et al.
Published: (2024)
by: Lin, Shihan, et al.
Published: (2024)
Context-Aware Hybrid Routing in Bluetooth Mesh Networks Using Multi-Model Machine Learning and AODV Fallback
by: Islam, Md Sajid, et al.
Published: (2025)
by: Islam, Md Sajid, et al.
Published: (2025)
Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane
by: Zhang, Yinchao, et al.
Published: (2025)
by: Zhang, Yinchao, et al.
Published: (2025)
LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems
by: Duan, Tianyang, et al.
Published: (2025)
by: Duan, Tianyang, et al.
Published: (2025)
IRS-Assisted Lossy Communications Under Correlated Rayleigh Fading: Outage Probability Analysis and Optimization
by: Li, Guanchang, et al.
Published: (2024)
by: Li, Guanchang, et al.
Published: (2024)
Data Driven Environmental Awareness Using Wireless Signals
by: Nasiri, Hossein, et al.
Published: (2024)
by: Nasiri, Hossein, et al.
Published: (2024)
Fast Heterogeneous Serving: Scalable Mixed-Scale LLM Allocation for SLO-Constrained Inference
by: Cheng, Jiaming, et al.
Published: (2026)
by: Cheng, Jiaming, et al.
Published: (2026)
Mobility-Aware Federated Self-supervised Learning in Vehicular Network
by: Gu, Xueying, et al.
Published: (2024)
by: Gu, Xueying, et al.
Published: (2024)
Conflict-Aware Client Selection for Multi-Server Federated Learning
by: Hong, Mingwei, et al.
Published: (2026)
by: Hong, Mingwei, et al.
Published: (2026)
Privacy-Aware Multi-Device Cooperative Edge Inference with Distributed Resource Bidding
by: Zhuang, Wenhao, et al.
Published: (2024)
by: Zhuang, Wenhao, et al.
Published: (2024)
Open World Learning Graph Convolution for Latency Estimation in Routing Networks
by: Jin, Yifei, et al.
Published: (2022)
by: Jin, Yifei, et al.
Published: (2022)
Latency-Distortion Tradeoffs in Communicating Classification Results over Noisy Channels
by: Teku, Noel, et al.
Published: (2024)
by: Teku, Noel, et al.
Published: (2024)
Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
by: Lin, Chungang, et al.
Published: (2025)
by: Lin, Chungang, et al.
Published: (2025)
Diffusion Models Meet Network Management: Improving Traffic Matrix Analysis with Diffusion-based Approach
by: Yuan, Xinyu, et al.
Published: (2024)
by: Yuan, Xinyu, et al.
Published: (2024)
Link-Aware Energy-Frugal Continual Learning for Fault Detection in IoT Networks
by: Frederiksen, Henrik C. M., et al.
Published: (2025)
by: Frederiksen, Henrik C. M., et al.
Published: (2025)
Intrusion Detection on Resource-Constrained IoT Devices with Hardware-Aware ML and DL
by: Diab, Ali, et al.
Published: (2025)
by: Diab, Ali, et al.
Published: (2025)
Interference-Aware Emergent Random Access Protocol for Downlink LEO Satellite Networks
by: Lim, Chang-Yong, et al.
Published: (2024)
by: Lim, Chang-Yong, et al.
Published: (2024)
Similar Items
-
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
by: Nan, Zhaojun, et al.
Published: (2025) -
Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
by: Han, Yunchu, et al.
Published: (2025) -
C-MASS: Combinatorial Mobility-Aware Sensor Scheduling for Collaborative Perception with Second-Order Topology Approximation
by: Jia, Yukuan, et al.
Published: (2024) -
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
by: Cheng, Zhipeng, et al.
Published: (2025) -
Velocity and Density-Aware RRI Analysis and Optimization for AoI Minimization in IoV SPS
by: Ji, Maoxin, et al.
Published: (2025)