CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Laukemann, Jan, Gruber, Thomas, Hager, Georg, Oryspayev, Dossay, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
von: Alappat, Christie, et al.
Veröffentlicht: (2023)
von: Alappat, Christie, et al.
Veröffentlicht: (2023)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024)
von: Yu, Luo, et al.
Veröffentlicht: (2024)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
Cloud Resource Allocation with Convex Optimization
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
von: Machado, Rafael Ravedutti Lucio, et al.
Veröffentlicht: (2025)
von: Machado, Rafael Ravedutti Lucio, et al.
Veröffentlicht: (2025)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Improved vectorization of OpenCV algorithms for RISC-V CPUs
von: Volokitin, V. D., et al.
Veröffentlicht: (2023)
von: Volokitin, V. D., et al.
Veröffentlicht: (2023)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
On the Partitioning of GPU Power among Multi-Instances
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
THAPI: Tracing Heterogeneous APIs
von: Bekele, Solomon, et al.
Veröffentlicht: (2025)
von: Bekele, Solomon, et al.
Veröffentlicht: (2025)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
von: Senk, Johanna, et al.
Veröffentlicht: (2025)
von: Senk, Johanna, et al.
Veröffentlicht: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024) -
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026) -
Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
von: Alappat, Christie, et al.
Veröffentlicht: (2023)