Dynamic Memory Management on GPUs with SYCL
Fuente:
arXiv
Saved in:
| Main Author: | Standish, Russell K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025)
by: Devarakonda, Aditya, et al.
Published: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
by: Yin, Jia, et al.
Published: (2025)
by: Yin, Jia, et al.
Published: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025)
by: Lappi, Oskar, et al.
Published: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
by: Bergmans, Kobe, et al.
Published: (2025)
by: Bergmans, Kobe, et al.
Published: (2025)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
A Morton-Type Space-Filling Curve for Pyramid Subdivision and Hybrid Adaptive Mesh Refinement
by: Knapp, David, et al.
Published: (2026)
by: Knapp, David, et al.
Published: (2026)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
by: Nicolaou, Nicolas, et al.
Published: (2026)
by: Nicolaou, Nicolas, et al.
Published: (2026)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Context Adaptive Cooperation
by: Albouy, Timothé, et al.
Published: (2023)
by: Albouy, Timothé, et al.
Published: (2023)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
by: Papp, Pál András, et al.
Published: (2024)
by: Papp, Pál András, et al.
Published: (2024)
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
by: Chowdhury, Tasnuva, et al.
Published: (2025)
by: Chowdhury, Tasnuva, et al.
Published: (2025)
Real-time chaotic video encryption based on multithreaded parallel confusion and diffusion
by: Jiang, Dong, et al.
Published: (2023)
by: Jiang, Dong, et al.
Published: (2023)
Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
by: Xiaoteng, et al.
Published: (2024)
by: Xiaoteng, et al.
Published: (2024)
Leveraging Multi-Instance GPUs through moldable task scheduling
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
GPU-Accelerated Algorithms for Process Mapping
by: Samoldekin, Petr, et al.
Published: (2025)
by: Samoldekin, Petr, et al.
Published: (2025)
Similar Items
-
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025) -
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024) -
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024) -
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024) -
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)