Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Grbic, Dragana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025)
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025)
PICO: Performance Insights for Collective Operations
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Integrating High Performance In-Memory Data Streaming and In-Situ Visualization in Hybrid MPI+OpenMP PIC MC Simulations Towards Exascale
von: Williams, Jeremy J., et al.
Veröffentlicht: (2025)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2026)
THAPI: Tracing Heterogeneous APIs
von: Bekele, Solomon, et al.
Veröffentlicht: (2025)
von: Bekele, Solomon, et al.
Veröffentlicht: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
A Performance Analysis of BFT Consensus for Blockchains
von: Chan, J. D., et al.
Veröffentlicht: (2024)
von: Chan, J. D., et al.
Veröffentlicht: (2024)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial
von: Thomasian, Alexander
Veröffentlicht: (2023)
von: Thomasian, Alexander
Veröffentlicht: (2023)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications
von: Godoy, William F., et al.
Veröffentlicht: (2025)
von: Godoy, William F., et al.
Veröffentlicht: (2025)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
eBPF-Based Instrumentation for Generalisable Diagnosis of Performance Degradation
von: Landau, Diogo, et al.
Veröffentlicht: (2025)
von: Landau, Diogo, et al.
Veröffentlicht: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Ähnliche Einträge
-
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024) -
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
von: Banegas-Luna, Antonio Jesús, et al.
Veröffentlicht: (2025) -
PICO: Performance Insights for Collective Operations
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025) -
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026) -
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)