Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brooks, Alex, Marshall, Philip, Ozog, David, Rahman, Md. Wasi-ur-, Stewart, Lawrence, Tom, Rithwik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
von: Frey, Matthias, et al.
Veröffentlicht: (2024)
von: Frey, Matthias, et al.
Veröffentlicht: (2024)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
von: Fridman, Yehonatan, et al.
Veröffentlicht: (2024)
von: Fridman, Yehonatan, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
Assessing Opportunities of SYCL for Biological Sequence Alignment on GPU-based Systems
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022)
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
von: Torres, L. A., et al.
Veröffentlicht: (2024)
von: Torres, L. A., et al.
Veröffentlicht: (2024)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
von: Costanzo, Manuel, et al.
Veröffentlicht: (2024)
von: Costanzo, Manuel, et al.
Veröffentlicht: (2024)
SYCL compute kernels for ExaHyPE
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
Rafture: Erasure-coded Raft with Post-Dissemination Pruning
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)
von: Tramm, John, et al.
Veröffentlicht: (2024)
Evaluating SYCL as a Unified Programming Model for Heterogeneous Systems
von: Marowka, Ami
Veröffentlicht: (2026)
von: Marowka, Ami
Veröffentlicht: (2026)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
Finding Nemo-Nemo: CFT DAG-based Consensus in the WAN
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
A GPU Accelerated Temporal Window-Based Random Walk Sampler
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
City-Scale Visibility Graph Analysis via GPU-Accelerated HyperBall
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
A Theoretical Framework for Graph-based Digital Twins for Supply Chain Management and Optimization
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2025)
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2025)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
Supporting Intel(r) SGX on Multi-Package Platforms
von: Johnson, Simon, et al.
Veröffentlicht: (2025)
von: Johnson, Simon, et al.
Veröffentlicht: (2025)
Gaining Cross-Platform Parallelism for HAL's Molecular Dynamics Package using SYCL
von: Skoblin, Viktor, et al.
Veröffentlicht: (2024)
von: Skoblin, Viktor, et al.
Veröffentlicht: (2024)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
von: Latt, Jonas, et al.
Veröffentlicht: (2025)
von: Latt, Jonas, et al.
Veröffentlicht: (2025)
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
von: Alekseenko, Andrey, et al.
Veröffentlicht: (2024)
Evaluation of Intel Max GPUs for CGYRO-based fusion simulations
von: Sfiligoi, Igor, et al.
Veröffentlicht: (2024)
von: Sfiligoi, Igor, et al.
Veröffentlicht: (2024)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
von: V., Karthik, et al.
Veröffentlicht: (2024)
von: V., Karthik, et al.
Veröffentlicht: (2024)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
GPU Sharing with Triples Mode
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
DuaLip-GPU Technical Report
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
von: Frey, Matthias, et al.
Veröffentlicht: (2024) -
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
von: Fridman, Yehonatan, et al.
Veröffentlicht: (2024) -
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022) -
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025) -
Assessing Opportunities of SYCL for Biological Sequence Alignment on GPU-based Systems
von: Costanzo, Manuel, et al.
Veröffentlicht: (2022)