Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Costanzo, Manuel, Rucci, Enzo, García-Sánchez, Carlos, Naiouf, Marcelo, Prieto-Matías, Manuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing Opportunities of SYCL for Biological Sequence Alignment on GPU-based Systems
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022)
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
Enhanced OpenMP Algorithm to Compute All-Pairs Shortest Path on x86 Architectures
von: Calderón, Sergio, et al.
Veröffentlicht: (2024)
von: Calderón, Sergio, et al.
Veröffentlicht: (2024)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)
von: Tramm, John, et al.
Veröffentlicht: (2024)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025)
von: Standish, Russell K.
Veröffentlicht: (2025)
HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs
von: Li, Yanliang, et al.
Veröffentlicht: (2025)
von: Li, Yanliang, et al.
Veröffentlicht: (2025)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
Portability Efficiency Approach for Calculating Performance Portability
von: Marowka, Ami
Veröffentlicht: (2024)
von: Marowka, Ami
Veröffentlicht: (2024)
Maple: A Multi-agent System for Portable Deep Learning across Clusters
von: Wu, Molang, et al.
Veröffentlicht: (2025)
von: Wu, Molang, et al.
Veröffentlicht: (2025)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
von: Zwaka, Linus
Veröffentlicht: (2024)
von: Zwaka, Linus
Veröffentlicht: (2024)
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
von: Brooks, Alex, et al.
Veröffentlicht: (2024)
von: Brooks, Alex, et al.
Veröffentlicht: (2024)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
Evaluating SYCL as a Unified Programming Model for Heterogeneous Systems
von: Marowka, Ami
Veröffentlicht: (2026)
von: Marowka, Ami
Veröffentlicht: (2026)
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
SYCL compute kernels for ExaHyPE
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
LAPIS: A Performance Portable, High Productivity Compiler Framework
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
HPDR: High-Performance Portable Scientific Data Reduction Framework
von: Chen, Jieyang, et al.
Veröffentlicht: (2025)
von: Chen, Jieyang, et al.
Veröffentlicht: (2025)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
von: Copik, Marcin, et al.
Veröffentlicht: (2025)
von: Copik, Marcin, et al.
Veröffentlicht: (2025)
Decentralized and Self-adaptive Core Maintenance on Temporal Graphs
von: Rucci, Davide, et al.
Veröffentlicht: (2025)
von: Rucci, Davide, et al.
Veröffentlicht: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
miniLB: A Performance Portability Study of Lattice-Boltzmann Simulations
von: Crisci, Luigi, et al.
Veröffentlicht: (2024)
von: Crisci, Luigi, et al.
Veröffentlicht: (2024)
Enabling performance portability of data-parallel OpenMP applications on asymmetric multicore processors
von: Saez, Juan Carlos, et al.
Veröffentlicht: (2024)
von: Saez, Juan Carlos, et al.
Veröffentlicht: (2024)
A dynamic parallel method for performance optimization on hybrid CPUs
von: Yu, Luo, et al.
Veröffentlicht: (2024)
von: Yu, Luo, et al.
Veröffentlicht: (2024)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
A Parallel and Distributed Rust Library for Core Decomposition on Large Graphs
von: Rucci, Davide, et al.
Veröffentlicht: (2025)
von: Rucci, Davide, et al.
Veröffentlicht: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
von: Sakib, Nazmus, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Assessing Opportunities of SYCL for Biological Sequence Alignment on GPU-based Systems
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022) -
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022) -
Enhanced OpenMP Algorithm to Compute All-Pairs Shortest Path on x86 Architectures
von: Calderón, Sergio, et al.
Veröffentlicht: (2024) -
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025) -
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)