PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Truong-Thanh, La, Hoang-Loc, Taherkordi, Amir, Eliassen, Frank, and, Phuong Hoai Ha, Guan, Peiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kernel-Level Energy-Efficient Neural Architecture Search for Tabular Dataset
by: La, Hoang-Loc, et al.
Published: (2025)
by: La, Hoang-Loc, et al.
Published: (2025)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
by: Jain, Rishabh, et al.
Published: (2024)
by: Jain, Rishabh, et al.
Published: (2024)
Fast Entropy Decoding for Sparse MVM on GPUs
by: Schätzle, Emil, et al.
Published: (2026)
by: Schätzle, Emil, et al.
Published: (2026)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
by: Provost, François, et al.
Published: (2025)
by: Provost, François, et al.
Published: (2025)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
by: Dagli, Ismet, et al.
Published: (2023)
by: Dagli, Ismet, et al.
Published: (2023)
A high-performance and portable implementation of the SISSO method for CPUs and GPUs
by: Eibl, Sebastian, et al.
Published: (2025)
by: Eibl, Sebastian, et al.
Published: (2025)
Automated PMC-based Power Modeling Methodology for Modern Mobile GPUs
by: Dash, Pranab, et al.
Published: (2024)
by: Dash, Pranab, et al.
Published: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023)
by: Kandel, Narayan, et al.
Published: (2023)
AFarePart: Accuracy-aware Fault-resilient Partitioner for DNN Edge Accelerators
by: Debnath, Mukta, et al.
Published: (2025)
by: Debnath, Mukta, et al.
Published: (2025)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
by: Li, Jianhui, et al.
Published: (2023)
by: Li, Jianhui, et al.
Published: (2023)
CarbonCP: Carbon-Aware DNN Partitioning with Conformal Prediction for Sustainable Edge Intelligence
by: Ke, Hongyu, et al.
Published: (2024)
by: Ke, Hongyu, et al.
Published: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
by: Zhang, Yongkang, et al.
Published: (2024)
by: Zhang, Yongkang, et al.
Published: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
PrETi: Predicting Execution Time in Early Stage with LLVM and Machine Learning
by: Xu, Risheng, et al.
Published: (2025)
by: Xu, Risheng, et al.
Published: (2025)
Opening the Black Box: Performance Estimation during Code Generation for GPUs
by: Ernst, Dominik, et al.
Published: (2021)
by: Ernst, Dominik, et al.
Published: (2021)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
by: Sapkas, M., et al.
Published: (2025)
by: Sapkas, M., et al.
Published: (2025)
Accurate and Scalable Many-Node Simulation
by: Eyerman, Stijn, et al.
Published: (2024)
by: Eyerman, Stijn, et al.
Published: (2024)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
by: Farooq, Muhammad Umar
Published: (2026)
by: Farooq, Muhammad Umar
Published: (2026)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
FRSZ2 for In-Register Block Compression Inside GMRES on GPUs
by: Grützmacher, Thomas, et al.
Published: (2024)
by: Grützmacher, Thomas, et al.
Published: (2024)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
Latency Based Tiling
by: Cashman, Jack
Published: (2025)
by: Cashman, Jack
Published: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
RAVE: RISC-V Analyzer of Vector Executions, a QEMU tracing plugin
by: Vizcaino, Pablo, et al.
Published: (2024)
by: Vizcaino, Pablo, et al.
Published: (2024)
ZERNIPAX: A Fast and Accurate Zernike Polynomial Calculator in Python
by: Elmacioglu, Yigit Gunsur, et al.
Published: (2024)
by: Elmacioglu, Yigit Gunsur, et al.
Published: (2024)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
by: Georganas, Evangelos, et al.
Published: (2025)
by: Georganas, Evangelos, et al.
Published: (2025)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
by: Liu, Youyuan, et al.
Published: (2024)
by: Liu, Youyuan, et al.
Published: (2024)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
Characterizing and Understanding HGNN Training on GPUs
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
Enhancing Tropical Cyclone Path Forecasting with an Improved Transformer Network
by: Van Thanh, Nguyen, et al.
Published: (2025)
by: Van Thanh, Nguyen, et al.
Published: (2025)
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
by: Pham, Nghiem Thanh, et al.
Published: (2025)
by: Pham, Nghiem Thanh, et al.
Published: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
Similar Items
-
Kernel-Level Energy-Efficient Neural Architecture Search for Tabular Dataset
by: La, Hoang-Loc, et al.
Published: (2025) -
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025) -
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
by: Jain, Rishabh, et al.
Published: (2024) -
Fast Entropy Decoding for Sparse MVM on GPUs
by: Schätzle, Emil, et al.
Published: (2026) -
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)