Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
Fuente:
arXiv
Salvato in:
| Autori principali: | de Castro, Manuel, andújar, Francisco J., Osorio, Roberto R., Carratalá-Sáez, Rocío, Llanos, Diego R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
di: Guidotti, Nicolas Lee
Pubblicazione: (2025)
di: Guidotti, Nicolas Lee
Pubblicazione: (2025)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
di: Thüring, Tim, et al.
Pubblicazione: (2026)
di: Thüring, Tim, et al.
Pubblicazione: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024)
di: Apanasevich, L., et al.
Pubblicazione: (2024)
Improving the Efficiency of OpenCL Kernels through Pipes
di: Zarch, Mostafa Eghbali, et al.
Pubblicazione: (2022)
di: Zarch, Mostafa Eghbali, et al.
Pubblicazione: (2022)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
di: Alekseenko, Andrey, et al.
Pubblicazione: (2024)
di: Alekseenko, Andrey, et al.
Pubblicazione: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
di: Pilliat, Emmanuel
Pubblicazione: (2026)
di: Pilliat, Emmanuel
Pubblicazione: (2026)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
di: Accordi, Gianmarco, et al.
Pubblicazione: (2025)
di: Accordi, Gianmarco, et al.
Pubblicazione: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
di: Panova, Elena, et al.
Pubblicazione: (2022)
di: Panova, Elena, et al.
Pubblicazione: (2022)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
di: Villalobos, Johansell, et al.
Pubblicazione: (2025)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
di: Costanzo, Manuel, et al.
Pubblicazione: (2024)
di: Costanzo, Manuel, et al.
Pubblicazione: (2024)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
di: Catalán, S., et al.
Pubblicazione: (2026)
di: Catalán, S., et al.
Pubblicazione: (2026)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
di: Sheng, Zhang, et al.
Pubblicazione: (2025)
di: Sheng, Zhang, et al.
Pubblicazione: (2025)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
di: Schneider, Timo, et al.
Pubblicazione: (2024)
di: Schneider, Timo, et al.
Pubblicazione: (2024)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
di: Rahimi, Ghazal, et al.
Pubblicazione: (2026)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
di: Pennino, Diego, et al.
Pubblicazione: (2024)
di: Pennino, Diego, et al.
Pubblicazione: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
di: Shan, Baodi, et al.
Pubblicazione: (2024)
di: Shan, Baodi, et al.
Pubblicazione: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
di: Lyu, Hanzheng, et al.
Pubblicazione: (2024)
di: Lyu, Hanzheng, et al.
Pubblicazione: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
di: Owen, Herbert, et al.
Pubblicazione: (2024)
di: Owen, Herbert, et al.
Pubblicazione: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
di: Wadhwa, Eashan, et al.
Pubblicazione: (2024)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2024)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
di: Williams, Jeremy J., et al.
Pubblicazione: (2024)
di: Williams, Jeremy J., et al.
Pubblicazione: (2024)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
di: Gonzalez-Escribano, Arturo, et al.
Pubblicazione: (2024)
di: Gonzalez-Escribano, Arturo, et al.
Pubblicazione: (2024)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
A Performance Analysis of BFT Consensus for Blockchains
di: Chan, J. D., et al.
Pubblicazione: (2024)
di: Chan, J. D., et al.
Pubblicazione: (2024)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
di: Wang, Zirui, et al.
Pubblicazione: (2026)
di: Wang, Zirui, et al.
Pubblicazione: (2026)
Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications
di: Godoy, William F., et al.
Pubblicazione: (2025)
di: Godoy, William F., et al.
Pubblicazione: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
di: Mao, Ying, et al.
Pubblicazione: (2020)
di: Mao, Ying, et al.
Pubblicazione: (2020)
Documenti analoghi
-
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
di: Guidotti, Nicolas Lee
Pubblicazione: (2025) -
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
di: Thüring, Tim, et al.
Pubblicazione: (2026) -
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
di: Apanasevich, L., et al.
Pubblicazione: (2024) -
Improving the Efficiency of OpenCL Kernels through Pipes
di: Zarch, Mostafa Eghbali, et al.
Pubblicazione: (2022) -
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)