Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sakib, Nazmus, Prabhu, Tarun, Santhi, Nandakishore, Shalf, John, Badawy, Abdel-Hameed A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024)
von: Yu, Luo, et al.
Veröffentlicht: (2024)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
von: Kozinov, Evgeny, et al.
Veröffentlicht: (2024)
von: Kozinov, Evgeny, et al.
Veröffentlicht: (2024)
Improved vectorization of OpenCV algorithms for RISC-V CPUs
von: Volokitin, V. D., et al.
Veröffentlicht: (2023)
von: Volokitin, V. D., et al.
Veröffentlicht: (2023)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025) -
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024) -
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024) -
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023) -
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)