On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
Fuente:
arXiv
Salvato in:
| Autori principali: | Machado, Rafael Ravedutti Lucio, Eitzinger, Jan, Hager, Georg, Wellein, Gerhard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
di: Laukemann, Jan, et al.
Pubblicazione: (2024)
di: Laukemann, Jan, et al.
Pubblicazione: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
di: Laukemann, Jan, et al.
Pubblicazione: (2023)
di: Laukemann, Jan, et al.
Pubblicazione: (2023)
Analytical Performance Estimation during Code Generation on Modern GPUs
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
Addressing Reproducibility Challenges in HPC with Continuous Integration
di: Hayot-Sasson, Valérie, et al.
Pubblicazione: (2025)
di: Hayot-Sasson, Valérie, et al.
Pubblicazione: (2025)
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
An Analysis of HPC and Edge Architectures in the Cloud
di: Santillan, Steven, et al.
Pubblicazione: (2025)
di: Santillan, Steven, et al.
Pubblicazione: (2025)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
di: Singh, Abhinav, et al.
Pubblicazione: (2023)
di: Singh, Abhinav, et al.
Pubblicazione: (2023)
Enabling mixed-precision in spectral element codes
di: Chen, Yanxiang, et al.
Pubblicazione: (2025)
di: Chen, Yanxiang, et al.
Pubblicazione: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
High-Performance Star-M SVD for Big Data Compression
di: Hussain, Md Taufique, et al.
Pubblicazione: (2026)
di: Hussain, Md Taufique, et al.
Pubblicazione: (2026)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
di: Bisbas, George, et al.
Pubblicazione: (2024)
di: Bisbas, George, et al.
Pubblicazione: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
di: Zhu, Honglin, et al.
Pubblicazione: (2026)
di: Zhu, Honglin, et al.
Pubblicazione: (2026)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
di: Havdiak, Mykhailo, et al.
Pubblicazione: (2024)
di: Havdiak, Mykhailo, et al.
Pubblicazione: (2024)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
di: Wang, Hansheng, et al.
Pubblicazione: (2025)
di: Wang, Hansheng, et al.
Pubblicazione: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
di: Carrica, Vicki, et al.
Pubblicazione: (2025)
di: Carrica, Vicki, et al.
Pubblicazione: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
di: Ham, David A., et al.
Pubblicazione: (2024)
di: Ham, David A., et al.
Pubblicazione: (2024)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
di: Merson, Jacob, et al.
Pubblicazione: (2023)
di: Merson, Jacob, et al.
Pubblicazione: (2023)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
di: Derlatka, Kacper, et al.
Pubblicazione: (2024)
di: Derlatka, Kacper, et al.
Pubblicazione: (2024)
SYCL compute kernels for ExaHyPE
di: Loi, Chung Ming, et al.
Pubblicazione: (2023)
di: Loi, Chung Ming, et al.
Pubblicazione: (2023)
Energy-aware operation of HPC systems in Germany
di: Suarez, Estela, et al.
Pubblicazione: (2024)
di: Suarez, Estela, et al.
Pubblicazione: (2024)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
di: Alappat, Christie, et al.
Pubblicazione: (2023)
di: Alappat, Christie, et al.
Pubblicazione: (2023)
Enabling mixed-precision with the help of tools: A Nekbone case study
di: Chen, Yanxiang, et al.
Pubblicazione: (2024)
di: Chen, Yanxiang, et al.
Pubblicazione: (2024)
Evaluation of POSIT Arithmetic with Accelerators
di: Nakasato, Naohito, et al.
Pubblicazione: (2024)
di: Nakasato, Naohito, et al.
Pubblicazione: (2024)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
di: Katagiri, Takahiro, et al.
Pubblicazione: (2024)
di: Katagiri, Takahiro, et al.
Pubblicazione: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
di: Bisbas, George, et al.
Pubblicazione: (2023)
di: Bisbas, George, et al.
Pubblicazione: (2023)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
di: Bernaschi, Massimo, et al.
Pubblicazione: (2025)
di: Bernaschi, Massimo, et al.
Pubblicazione: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
di: Woller, Fabian, et al.
Pubblicazione: (2025)
di: Woller, Fabian, et al.
Pubblicazione: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
di: Katagiri, Takahiro, et al.
Pubblicazione: (2024)
di: Katagiri, Takahiro, et al.
Pubblicazione: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
Performance measurements of modern Fortran MPI applications with Score-P
di: Corbin, Gregor
Pubblicazione: (2025)
di: Corbin, Gregor
Pubblicazione: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
di: Panova, Elena, et al.
Pubblicazione: (2022)
di: Panova, Elena, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2024) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2026) -
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
di: Laukemann, Jan, et al.
Pubblicazione: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025) -
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
di: Laukemann, Jan, et al.
Pubblicazione: (2023)