Saved in:
| Main Authors: | Mayr, Martin, Wind, Sebastian, Schröder, Lukas, Hager, Georg, Köstler, Harald, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.16164 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025)
by: Afzal, Ayesha, et al.
Published: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200
by: Ujeniya, Aditya, et al.
Published: (2026)
by: Ujeniya, Aditya, et al.
Published: (2026)
Opening the Black Box: Performance Estimation during Code Generation for GPUs
by: Ernst, Dominik, et al.
Published: (2021)
by: Ernst, Dominik, et al.
Published: (2021)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023)
by: Laukemann, Jan, et al.
Published: (2023)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)
by: Alt, Christoph, et al.
Published: (2024)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
by: Suffa, Philipp, et al.
Published: (2024)
by: Suffa, Philipp, et al.
Published: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024)
by: Owen, Herbert, et al.
Published: (2024)
SteuerLLM: Local specialized large language model for German tax law analysis
by: Wind, Sebastian, et al.
Published: (2026)
by: Wind, Sebastian, et al.
Published: (2026)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
by: Liu, Youyuan, et al.
Published: (2024)
by: Liu, Youyuan, et al.
Published: (2024)
Large-scale Multigrid with Adaptive Galerkin Coarsening
by: Böhm, Fabian, et al.
Published: (2025)
by: Böhm, Fabian, et al.
Published: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
by: Werner, Elias, et al.
Published: (2023)
by: Werner, Elias, et al.
Published: (2023)
OSCAR-P and aMLLibrary: Profiling and Predicting the Performance of FaaS-based Applications in Computing Continua
by: Sala, Roberto, et al.
Published: (2024)
by: Sala, Roberto, et al.
Published: (2024)
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)
by: Kumar, Harsh, et al.
Published: (2024)
Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
by: Ritter, Fabian, et al.
Published: (2024)
by: Ritter, Fabian, et al.
Published: (2024)
More than "Just" Music: Four Performative Topoi, the Phish Phenomenon, and the Power of Music in/and Performance
by: Jnan Blau
Published: (2009)
by: Jnan Blau
Published: (2009)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
by: Salaria, Shweta, et al.
Published: (2025)
by: Salaria, Shweta, et al.
Published: (2025)
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
Resource Allocation Influence on Application Performance in Sliced Testbeds
by: Moreira, Rodrigo, et al.
Published: (2024)
by: Moreira, Rodrigo, et al.
Published: (2024)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
by: Zheng, Yusheng, et al.
Published: (2026)
by: Zheng, Yusheng, et al.
Published: (2026)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
by: Mazzola, Sergio, et al.
Published: (2025)
by: Mazzola, Sergio, et al.
Published: (2025)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
by: Yu, Dingcui, et al.
Published: (2025)
by: Yu, Dingcui, et al.
Published: (2025)
A Priori Loop Nest Normalization: Automatic Loop Scheduling in Complex Applications
by: Trümper, Lukas, et al.
Published: (2024)
by: Trümper, Lukas, et al.
Published: (2024)
Impact of Generative AI (Large Language Models) on the PRA model construction and maintenance, observations
by: Rychkov, Valentin, et al.
Published: (2024)
by: Rychkov, Valentin, et al.
Published: (2024)
An Analysis of Performance Bottlenecks in MRI Pre-Processing
by: Dugré, Mathieu, et al.
Published: (2024)
by: Dugré, Mathieu, et al.
Published: (2024)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
Noise Injection for__Performance Bottleneck Analysis
by: Delval, Aurélien, et al.
Published: (2025)
by: Delval, Aurélien, et al.
Published: (2025)
Enabling Heterogeneous Performance Analysis for Scientific Workloads
by: Graczyk, Maksymilian, et al.
Published: (2025)
by: Graczyk, Maksymilian, et al.
Published: (2025)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
by: Wang, Yulei, et al.
Published: (2025)
by: Wang, Yulei, et al.
Published: (2025)
AMD MI300X GPU Performance Analysis
by: Ambati, Chandrish, et al.
Published: (2025)
by: Ambati, Chandrish, et al.
Published: (2025)
AI Load Dynamics--A Power Electronics Perspective
by: Li, Yuzhuo, et al.
Published: (2025)
by: Li, Yuzhuo, et al.
Published: (2025)
Similar Items
-
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026) -
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024) -
Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200
by: Ujeniya, Aditya, et al.
Published: (2026)