Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
Fuente:
arXiv
Salvato in:
| Autori principali: | Laukemann, Jan, Hager, Georg, Wellein, Gerhard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
di: Laukemann, Jan, et al.
Pubblicazione: (2023)
di: Laukemann, Jan, et al.
Pubblicazione: (2023)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
di: Afzal, Ayesha, et al.
Pubblicazione: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
di: Afzal, Ayesha, et al.
Pubblicazione: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
di: Alappat, Christie, et al.
Pubblicazione: (2023)
di: Alappat, Christie, et al.
Pubblicazione: (2023)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
di: Dutilleul, Alban, et al.
Pubblicazione: (2024)
di: Dutilleul, Alban, et al.
Pubblicazione: (2024)
A dynamic parallel method for performance optimization on hybrid CPUs
di: Yu, Luo, et al.
Pubblicazione: (2024)
di: Yu, Luo, et al.
Pubblicazione: (2024)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
di: Wolfson-Pou, Jordi, et al.
Pubblicazione: (2025)
di: Wolfson-Pou, Jordi, et al.
Pubblicazione: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
di: Sakib, Nazmus, et al.
Pubblicazione: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
di: Accordi, Gianmarco, et al.
Pubblicazione: (2025)
di: Accordi, Gianmarco, et al.
Pubblicazione: (2025)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
di: Owen, Herbert, et al.
Pubblicazione: (2024)
di: Owen, Herbert, et al.
Pubblicazione: (2024)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
di: Muhammad, Said, et al.
Pubblicazione: (2025)
di: Muhammad, Said, et al.
Pubblicazione: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
Analytical Performance Estimation during Code Generation on Modern GPUs
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
di: Helal, Ahmed E., et al.
Pubblicazione: (2025)
di: Helal, Ahmed E., et al.
Pubblicazione: (2025)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
di: Machado, Rafael Ravedutti Lucio, et al.
Pubblicazione: (2025)
di: Machado, Rafael Ravedutti Lucio, et al.
Pubblicazione: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
di: Panova, Elena, et al.
Pubblicazione: (2022)
di: Panova, Elena, et al.
Pubblicazione: (2022)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
di: Ibeid, Huda, et al.
Pubblicazione: (2025)
di: Ibeid, Huda, et al.
Pubblicazione: (2025)
Improved vectorization of OpenCV algorithms for RISC-V CPUs
di: Volokitin, V. D., et al.
Pubblicazione: (2023)
di: Volokitin, V. D., et al.
Pubblicazione: (2023)
Performance and scaling of the LFRic weather and climate model on different generations of HPE Cray EX supercomputers
di: Bull, J. Mark, et al.
Pubblicazione: (2024)
di: Bull, J. Mark, et al.
Pubblicazione: (2024)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
di: Senk, Johanna, et al.
Pubblicazione: (2025)
di: Senk, Johanna, et al.
Pubblicazione: (2025)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
di: Laukemann, Jan, et al.
Pubblicazione: (2024)
di: Laukemann, Jan, et al.
Pubblicazione: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
Staging Blocked Evaluation over Structured Sparse Matrices
di: Das, Pratyush, et al.
Pubblicazione: (2024)
di: Das, Pratyush, et al.
Pubblicazione: (2024)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
Stabl: Blockchain Fault Tolerance
di: Gramoli, Vincent, et al.
Pubblicazione: (2024)
di: Gramoli, Vincent, et al.
Pubblicazione: (2024)
EfiMon: A Process Analyser for Granular Power Consumption Prediction
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
di: Suffa, Philipp, et al.
Pubblicazione: (2024)
di: Suffa, Philipp, et al.
Pubblicazione: (2024)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
di: Pennino, Diego, et al.
Pubblicazione: (2024)
di: Pennino, Diego, et al.
Pubblicazione: (2024)
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
di: Solórzano, Ana Luisa Veroneze, et al.
Pubblicazione: (2024)
di: Solórzano, Ana Luisa Veroneze, et al.
Pubblicazione: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
di: Lurati, Milo, et al.
Pubblicazione: (2024)
di: Lurati, Milo, et al.
Pubblicazione: (2024)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
di: Chien, Steven W. D., et al.
Pubblicazione: (2024)
di: Chien, Steven W. D., et al.
Pubblicazione: (2024)
Documenti analoghi
-
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
di: Laukemann, Jan, et al.
Pubblicazione: (2023) -
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
di: Afzal, Ayesha, et al.
Pubblicazione: (2026) -
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)