Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jolly, Dessole, Monica, Varbanescu, Ana Lucia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
by: Apanasevich, L., et al.
Published: (2024)
by: Apanasevich, L., et al.
Published: (2024)
Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies
by: Brakel, Felix, et al.
Published: (2024)
by: Brakel, Felix, et al.
Published: (2024)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
by: Cosenza, Biagio, et al.
Published: (2025)
by: Cosenza, Biagio, et al.
Published: (2025)
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
by: Brooks, Alex, et al.
Published: (2024)
by: Brooks, Alex, et al.
Published: (2024)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)
by: Costanzo, Manuel, et al.
Published: (2024)
SkimROOT: Accelerating LHC Data Filtering with Near-Storage Processing
by: Batsoyol, Narangerelt, et al.
Published: (2025)
by: Batsoyol, Narangerelt, et al.
Published: (2025)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
Parallel Gaussian process with kernel approximation in CUDA
by: Carminati, Davide
Published: (2024)
by: Carminati, Davide
Published: (2024)
SYCL compute kernels for ExaHyPE
by: Loi, Chung Ming, et al.
Published: (2023)
by: Loi, Chung Ming, et al.
Published: (2023)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
by: Thüring, Tim, et al.
Published: (2026)
by: Thüring, Tim, et al.
Published: (2026)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
Evaluating SYCL as a Unified Programming Model for Heterogeneous Systems
by: Marowka, Ami
Published: (2026)
by: Marowka, Ami
Published: (2026)
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
by: ALHafez, Nizar, et al.
Published: (2025)
by: ALHafez, Nizar, et al.
Published: (2025)
Assessing Opportunities of SYCL for Biological Sequence Alignment on GPU-based Systems
by: Costanzo, Manuel, et al.
Published: (2022)
by: Costanzo, Manuel, et al.
Published: (2022)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
by: Tolotti, Emiliano, et al.
Published: (2024)
by: Tolotti, Emiliano, et al.
Published: (2024)
Lessons Learned on the Path to Guaranteeing the Error Bound in Lossy Quantizers
by: Fallin, Alex, et al.
Published: (2024)
by: Fallin, Alex, et al.
Published: (2024)
Lessons Learned from Building Edge Software System Testbeds
by: Pfandzelter, Tobias, et al.
Published: (2024)
by: Pfandzelter, Tobias, et al.
Published: (2024)
Tutoring LLM into a Better CUDA Optimizer
by: Brabec, Matyáš, et al.
Published: (2025)
by: Brabec, Matyáš, et al.
Published: (2025)
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
by: Li, Xiaoya, et al.
Published: (2025)
by: Li, Xiaoya, et al.
Published: (2025)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
by: Torres, L. A., et al.
Published: (2024)
by: Torres, L. A., et al.
Published: (2024)
cuConv: A CUDA Implementation of Convolution for CNN Inference
by: Jordà, Marc, et al.
Published: (2021)
by: Jordà, Marc, et al.
Published: (2021)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Gaining Cross-Platform Parallelism for HAL's Molecular Dynamics Package using SYCL
by: Skoblin, Viktor, et al.
Published: (2024)
by: Skoblin, Viktor, et al.
Published: (2024)
Seamless Transitions: A Comprehensive Review of Live Migration Technologies
by: Attar-Khorasani, Sima, et al.
Published: (2025)
by: Attar-Khorasani, Sima, et al.
Published: (2025)
Buffer Centering for bittide Synchronization via Frame Rotation
by: Lall, Sanjay, et al.
Published: (2025)
by: Lall, Sanjay, et al.
Published: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
by: Chen, Fahao, et al.
Published: (2024)
by: Chen, Fahao, et al.
Published: (2024)
HopGNN: Boosting Distributed GNN Training Efficiency via Feature-Centric Model Migration
by: Chen, Weijian, et al.
Published: (2024)
by: Chen, Weijian, et al.
Published: (2024)
Modeling Distributed Computing Infrastructures for HEP Applications
by: Horzela, Maximilian, et al.
Published: (2024)
by: Horzela, Maximilian, et al.
Published: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
by: Goto, Kazushige, et al.
Published: (2026)
by: Goto, Kazushige, et al.
Published: (2026)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
CUDA Kernel Optimization and Counter-Free Performance Analysis for Depthwise Convolution in Cloud Environments
by: Babak, Huriyeh, et al.
Published: (2026)
by: Babak, Huriyeh, et al.
Published: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
CausalMesh: A Formally Verified Causally Consistent Distributed Cache with Support for Client Migration
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
Similar Items
-
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
by: Apanasevich, L., et al.
Published: (2024) -
Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies
by: Brakel, Felix, et al.
Published: (2024) -
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
by: Cosenza, Biagio, et al.
Published: (2025) -
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
by: Brooks, Alex, et al.
Published: (2024) -
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)