PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Le, Truong-Thanh, La, Hoang-Loc, Taherkordi, Amir, Eliassen, Frank, and, Phuong Hoai Ha, Guan, Peiyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Kernel-Level Energy-Efficient Neural Architecture Search for Tabular Dataset
di: La, Hoang-Loc, et al.
Pubblicazione: (2025)
di: La, Hoang-Loc, et al.
Pubblicazione: (2025)
Performance of Confidential Computing GPUs
di: Ibarra, Antonio Martínez, et al.
Pubblicazione: (2025)
di: Ibarra, Antonio Martínez, et al.
Pubblicazione: (2025)
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
di: Jain, Rishabh, et al.
Pubblicazione: (2024)
di: Jain, Rishabh, et al.
Pubblicazione: (2024)
Fast Entropy Decoding for Sparse MVM on GPUs
di: Schätzle, Emil, et al.
Pubblicazione: (2026)
di: Schätzle, Emil, et al.
Pubblicazione: (2026)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2025)
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2025)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
di: Provost, François, et al.
Pubblicazione: (2025)
di: Provost, François, et al.
Pubblicazione: (2025)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
di: Dagli, Ismet, et al.
Pubblicazione: (2023)
di: Dagli, Ismet, et al.
Pubblicazione: (2023)
A high-performance and portable implementation of the SISSO method for CPUs and GPUs
di: Eibl, Sebastian, et al.
Pubblicazione: (2025)
di: Eibl, Sebastian, et al.
Pubblicazione: (2025)
Automated PMC-based Power Modeling Methodology for Modern Mobile GPUs
di: Dash, Pranab, et al.
Pubblicazione: (2024)
di: Dash, Pranab, et al.
Pubblicazione: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
di: Kandel, Narayan, et al.
Pubblicazione: (2023)
di: Kandel, Narayan, et al.
Pubblicazione: (2023)
AFarePart: Accuracy-aware Fault-resilient Partitioner for DNN Edge Accelerators
di: Debnath, Mukta, et al.
Pubblicazione: (2025)
di: Debnath, Mukta, et al.
Pubblicazione: (2025)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
CarbonCP: Carbon-Aware DNN Partitioning with Conformal Prediction for Sustainable Edge Intelligence
di: Ke, Hongyu, et al.
Pubblicazione: (2024)
di: Ke, Hongyu, et al.
Pubblicazione: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
di: Zhang, Yongkang, et al.
Pubblicazione: (2024)
di: Zhang, Yongkang, et al.
Pubblicazione: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
di: Hendria, Willy Fitra
Pubblicazione: (2026)
di: Hendria, Willy Fitra
Pubblicazione: (2026)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
di: Shen, Siyuan, et al.
Pubblicazione: (2025)
di: Shen, Siyuan, et al.
Pubblicazione: (2025)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
di: Ahmadvand, Hossein, et al.
Pubblicazione: (2025)
di: Ahmadvand, Hossein, et al.
Pubblicazione: (2025)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
di: Zhou, Fang, et al.
Pubblicazione: (2026)
di: Zhou, Fang, et al.
Pubblicazione: (2026)
How to Rent GPUs on a Budget
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
PrETi: Predicting Execution Time in Early Stage with LLVM and Machine Learning
di: Xu, Risheng, et al.
Pubblicazione: (2025)
di: Xu, Risheng, et al.
Pubblicazione: (2025)
Opening the Black Box: Performance Estimation during Code Generation for GPUs
di: Ernst, Dominik, et al.
Pubblicazione: (2021)
di: Ernst, Dominik, et al.
Pubblicazione: (2021)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
di: Sapkas, M., et al.
Pubblicazione: (2025)
di: Sapkas, M., et al.
Pubblicazione: (2025)
Accurate and Scalable Many-Node Simulation
di: Eyerman, Stijn, et al.
Pubblicazione: (2024)
di: Eyerman, Stijn, et al.
Pubblicazione: (2024)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
di: Farooq, Muhammad Umar
Pubblicazione: (2026)
di: Farooq, Muhammad Umar
Pubblicazione: (2026)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
di: Liu, Jiahui, et al.
Pubblicazione: (2024)
di: Liu, Jiahui, et al.
Pubblicazione: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
di: Liu, Yi
Pubblicazione: (2026)
di: Liu, Yi
Pubblicazione: (2026)
FRSZ2 for In-Register Block Compression Inside GMRES on GPUs
di: Grützmacher, Thomas, et al.
Pubblicazione: (2024)
di: Grützmacher, Thomas, et al.
Pubblicazione: (2024)
An Experimental Study of Low-Latency Video Streaming over 5G
di: Khan, Imran, et al.
Pubblicazione: (2024)
di: Khan, Imran, et al.
Pubblicazione: (2024)
Latency Based Tiling
di: Cashman, Jack
Pubblicazione: (2025)
di: Cashman, Jack
Pubblicazione: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
RAVE: RISC-V Analyzer of Vector Executions, a QEMU tracing plugin
di: Vizcaino, Pablo, et al.
Pubblicazione: (2024)
di: Vizcaino, Pablo, et al.
Pubblicazione: (2024)
ZERNIPAX: A Fast and Accurate Zernike Polynomial Calculator in Python
di: Elmacioglu, Yigit Gunsur, et al.
Pubblicazione: (2024)
di: Elmacioglu, Yigit Gunsur, et al.
Pubblicazione: (2024)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
di: Liu, Youyuan, et al.
Pubblicazione: (2024)
di: Liu, Youyuan, et al.
Pubblicazione: (2024)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
Characterizing and Understanding HGNN Training on GPUs
di: Han, Dengke, et al.
Pubblicazione: (2024)
di: Han, Dengke, et al.
Pubblicazione: (2024)
Enhancing Tropical Cyclone Path Forecasting with an Improved Transformer Network
di: Van Thanh, Nguyen, et al.
Pubblicazione: (2025)
di: Van Thanh, Nguyen, et al.
Pubblicazione: (2025)
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
di: Pham, Nghiem Thanh, et al.
Pubblicazione: (2025)
di: Pham, Nghiem Thanh, et al.
Pubblicazione: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
di: Ma, Ming, et al.
Pubblicazione: (2025)
di: Ma, Ming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Kernel-Level Energy-Efficient Neural Architecture Search for Tabular Dataset
di: La, Hoang-Loc, et al.
Pubblicazione: (2025) -
Performance of Confidential Computing GPUs
di: Ibarra, Antonio Martínez, et al.
Pubblicazione: (2025) -
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
di: Jain, Rishabh, et al.
Pubblicazione: (2024) -
Fast Entropy Decoding for Sparse MVM on GPUs
di: Schätzle, Emil, et al.
Pubblicazione: (2026) -
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2025)