Opening the Black Box: Performance Estimation during Code Generation for GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Ernst, Dominik, Hager, Georg, Holzer, Markus, Knorr, Matthias, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025)
by: Afzal, Ayesha, et al.
Published: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200
by: Ujeniya, Aditya, et al.
Published: (2026)
by: Ujeniya, Aditya, et al.
Published: (2026)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
by: Mayr, Martin, et al.
Published: (2026)
by: Mayr, Martin, et al.
Published: (2026)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024)
by: Owen, Herbert, et al.
Published: (2024)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023)
by: Laukemann, Jan, et al.
Published: (2023)
Decomposing Docker Container Startup Performance: A Three-Tier Measurement Study on Heterogeneous Infrastructure
by: Khan, Shamsher
Published: (2026)
by: Khan, Shamsher
Published: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
TurboMem: High-Performance Lock-Free Memory Pool with Transparent Huge Page Auto-Merging for DPDK
by: Yang, Junyi
Published: (2026)
by: Yang, Junyi
Published: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Stochastic Network Calculus with Localized Application of Martingales
by: Bouillard, Anne
Published: (2022)
by: Bouillard, Anne
Published: (2022)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Detection of Performance Changes in MooBench Results Using Nyrkiö on GitHub Actions
by: Yang, Shinhyung, et al.
Published: (2025)
by: Yang, Shinhyung, et al.
Published: (2025)
Performance Analysis of OpenVPN on a Consumer Grade Router
by: Hall, Michael J.
Published: (2025)
by: Hall, Michael J.
Published: (2025)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
A performance analysis of VM-based Trusted Execution Environments for Confidential Federated Learning
by: Casella, Bruno
Published: (2025)
by: Casella, Bruno
Published: (2025)
Photonic Fabric Platform for AI Accelerators
by: Ding, Jing, et al.
Published: (2025)
by: Ding, Jing, et al.
Published: (2025)
Finite-Time Behavior of Erlang-C Model: Mixing Time, Mean Queue Length and Tail Bounds
by: Nguyen, Hoang Huy, et al.
Published: (2025)
by: Nguyen, Hoang Huy, et al.
Published: (2025)
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
by: Li, Haorui, et al.
Published: (2026)
by: Li, Haorui, et al.
Published: (2026)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Glass-Box Analysis for Computer Systems: Transparency Index, Shapley Attribution, and Markov Models of Branch Prediction
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
Boosting Cross-Architectural Emulation Performance by Foregoing the Intermediate Representation Model
by: Parker, Amy Iris
Published: (2025)
by: Parker, Amy Iris
Published: (2025)
Fast Entropy Decoding for Sparse MVM on GPUs
by: Schätzle, Emil, et al.
Published: (2026)
by: Schätzle, Emil, et al.
Published: (2026)
A high-performance and portable implementation of the SISSO method for CPUs and GPUs
by: Eibl, Sebastian, et al.
Published: (2025)
by: Eibl, Sebastian, et al.
Published: (2025)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024)
by: Mukherjee, Dipayan, et al.
Published: (2024)
A note on integrating products of linear forms over the unit simplex
by: Casale, Giuliano
Published: (2017)
by: Casale, Giuliano
Published: (2017)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
by: Kharma, Sami, et al.
Published: (2026)
by: Kharma, Sami, et al.
Published: (2026)
KeyMemRT Compiler and Runtime: Unlocking Memory-Scalable FHE
by: Ünay, Eymen, et al.
Published: (2026)
by: Ünay, Eymen, et al.
Published: (2026)
Cost-effective and performant virtual WANs with CORNIFER
by: Anjali, et al.
Published: (2024)
by: Anjali, et al.
Published: (2024)
On the Power Saving in High-Speed Ethernet-based Networks for Supercomputers and Data Centers
by: de la Rosa, Miguel Sánchez, et al.
Published: (2025)
by: de la Rosa, Miguel Sánchez, et al.
Published: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
by: Ilager, Shashikant, et al.
Published: (2025)
by: Ilager, Shashikant, et al.
Published: (2025)
Predictive Modeling of I/O Performance for Machine Learning Training Pipelines: A Data-Driven Approach to Storage Optimization
by: Prabhakar, Karthik, et al.
Published: (2025)
by: Prabhakar, Karthik, et al.
Published: (2025)
On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II
by: Qiu, Jiaming, et al.
Published: (2026)
by: Qiu, Jiaming, et al.
Published: (2026)
Arm DynamIQ Shared Unit and Real-Time: An Empirical Evaluation
by: Pradhan, Ashutosh, et al.
Published: (2025)
by: Pradhan, Ashutosh, et al.
Published: (2025)
Similar Items
-
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026) -
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024) -
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)