Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Owen, Herbert, Ernst, Dominik, Gruber, Thomas, Lemkuhl, Oriol, Houzeaux, Guillaume, Gasparino, Lucas, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Opening the Black Box: Performance Estimation during Code Generation for GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2021)
von: Ernst, Dominik, et al.
Veröffentlicht: (2021)
Alya towards Exascale: Algorithmic Scalability using PSCToolkit
von: Owen, Herbert, et al.
Veröffentlicht: (2022)
von: Owen, Herbert, et al.
Veröffentlicht: (2022)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024)
Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200
von: Ujeniya, Aditya, et al.
Veröffentlicht: (2026)
von: Ujeniya, Aditya, et al.
Veröffentlicht: (2026)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
von: Mayr, Martin, et al.
Veröffentlicht: (2026)
von: Mayr, Martin, et al.
Veröffentlicht: (2026)
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
Performance of Confidential Computing GPUs
von: Ibarra, Antonio Martínez, et al.
Veröffentlicht: (2025)
von: Ibarra, Antonio Martínez, et al.
Veröffentlicht: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Fast Entropy Decoding for Sparse MVM on GPUs
von: Schätzle, Emil, et al.
Veröffentlicht: (2026)
von: Schätzle, Emil, et al.
Veröffentlicht: (2026)
Updates on the Low-Level Abstraction of Memory Access
von: Gruber, Bernhard Manfred
Veröffentlicht: (2023)
von: Gruber, Bernhard Manfred
Veröffentlicht: (2023)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2025)
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
Cache Blocking for Flux Reconstruction: Extension to Navier-Stokes Equations and Anti-aliasing
von: Akkurt, Semih, et al.
Veröffentlicht: (2023)
von: Akkurt, Semih, et al.
Veröffentlicht: (2023)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
von: Grbic, Dragana
Veröffentlicht: (2026)
von: Grbic, Dragana
Veröffentlicht: (2026)
A high-performance and portable implementation of the SISSO method for CPUs and GPUs
von: Eibl, Sebastian, et al.
Veröffentlicht: (2025)
von: Eibl, Sebastian, et al.
Veröffentlicht: (2025)
Automated PMC-based Power Modeling Methodology for Modern Mobile GPUs
von: Dash, Pranab, et al.
Veröffentlicht: (2024)
von: Dash, Pranab, et al.
Veröffentlicht: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
von: Kandel, Narayan, et al.
Veröffentlicht: (2023)
von: Kandel, Narayan, et al.
Veröffentlicht: (2023)
OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs
von: Wilfong, Benjamin, et al.
Veröffentlicht: (2024)
von: Wilfong, Benjamin, et al.
Veröffentlicht: (2024)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
von: Le, Truong-Thanh, et al.
Veröffentlicht: (2026)
von: Le, Truong-Thanh, et al.
Veröffentlicht: (2026)
Uncertainty Quantification as a Complementary Latent Health Indicator for Remaining Useful Life Prediction on Turbofan Engines
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
How to Rent GPUs on a Budget
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
von: Liu, Jiahui, et al.
Veröffentlicht: (2024)
von: Liu, Jiahui, et al.
Veröffentlicht: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
von: Liu, Yi
Veröffentlicht: (2026)
von: Liu, Yi
Veröffentlicht: (2026)
Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
FRSZ2 for In-Register Block Compression Inside GMRES on GPUs
von: Grützmacher, Thomas, et al.
Veröffentlicht: (2024)
von: Grützmacher, Thomas, et al.
Veröffentlicht: (2024)
Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications
von: Godoy, William F., et al.
Veröffentlicht: (2025)
von: Godoy, William F., et al.
Veröffentlicht: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Wasure: A Modular Toolkit for Comprehensive WebAssembly Benchmarking
von: Carissimi, Riccardo, et al.
Veröffentlicht: (2026)
von: Carissimi, Riccardo, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
Accelerating AI Performance using Anderson Extrapolation on GPUs
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Opening the Black Box: Performance Estimation during Code Generation for GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2021) -
Alya towards Exascale: Algorithmic Scalability using PSCToolkit
von: Owen, Herbert, et al.
Veröffentlicht: (2022) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025) -
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023) -
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
von: Williams, Jeremy J., et al.
Veröffentlicht: (2024)