Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ibeid, Huda, Narayana, Vikram, Kim, Jeongnim, Nguyen, Anthony, Morozov, Vitali, Luo, Ye |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
von: Jiang, Wenqi
Veröffentlicht: (2025)
von: Jiang, Wenqi
Veröffentlicht: (2025)
Scaling MPI Applications on Aurora
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024)
von: Yu, Luo, et al.
Veröffentlicht: (2024)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023) -
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026) -
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023) -
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)