Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ekelund, Jonah, Markidis, Stefano, Peng, Ivy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
por: Schieffer, Gabin, et al.
Publicado: (2025)
por: Schieffer, Gabin, et al.
Publicado: (2025)
QPU Micro-Kernels for Stencil Computation
por: Markidis, Stefano, et al.
Publicado: (2025)
por: Markidis, Stefano, et al.
Publicado: (2025)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
por: Williams, Jeremy J., et al.
Publicado: (2023)
por: Williams, Jeremy J., et al.
Publicado: (2023)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
Characterizing the Performance of the Implicit Massively Parallel Particle-in-Cell iPIC3D Code
por: Williams, Jeremy J., et al.
Publicado: (2024)
por: Williams, Jeremy J., et al.
Publicado: (2024)
Enabling AI Deep Potentials for Ab Initio-quality Molecular Dynamics Simulations in GROMACS
por: Hu, Andong, et al.
Publicado: (2026)
por: Hu, Andong, et al.
Publicado: (2026)
What is Quantum Parallelism, Anyhow?
por: Markidis, Stefano
Publicado: (2024)
por: Markidis, Stefano
Publicado: (2024)
Krylov Solvers for Interior Point Methods with Applications in Radiation Therapy and Support Vector Machines
por: Liu, Felix, et al.
Publicado: (2023)
por: Liu, Felix, et al.
Publicado: (2023)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
por: Song, Boyang
Publicado: (2025)
por: Song, Boyang
Publicado: (2025)
Making Room for AI: Multi-GPU Molecular Dynamics with Deep Potentials in GROMACS
por: Pennati, Luca, et al.
Publicado: (2026)
por: Pennati, Luca, et al.
Publicado: (2026)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
por: Shi, Ruimin, et al.
Publicado: (2025)
por: Shi, Ruimin, et al.
Publicado: (2025)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
por: Zwaka, Linus
Publicado: (2024)
por: Zwaka, Linus
Publicado: (2024)
CUDA Kernel Optimization and Counter-Free Performance Analysis for Depthwise Convolution in Cloud Environments
por: Babak, Huriyeh, et al.
Publicado: (2026)
por: Babak, Huriyeh, et al.
Publicado: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
por: Bellavita, Julian, et al.
Publicado: (2025)
por: Bellavita, Julian, et al.
Publicado: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
por: Peng, Ivy, et al.
Publicado: (2024)
por: Peng, Ivy, et al.
Publicado: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
por: Samsi, Siddharth, et al.
Publicado: (2025)
por: Samsi, Siddharth, et al.
Publicado: (2025)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
por: Li, Aiying, et al.
Publicado: (2026)
por: Li, Aiying, et al.
Publicado: (2026)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
por: Wahlgren, Jacob, et al.
Publicado: (2024)
por: Wahlgren, Jacob, et al.
Publicado: (2024)
Parallel Gaussian process with kernel approximation in CUDA
por: Carminati, Davide
Publicado: (2024)
por: Carminati, Davide
Publicado: (2024)
Analytical Performance Estimation during Code Generation on Modern GPUs
por: Ernst, Dominik, et al.
Publicado: (2022)
por: Ernst, Dominik, et al.
Publicado: (2022)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
por: Shi, Ruimin, et al.
Publicado: (2026)
por: Shi, Ruimin, et al.
Publicado: (2026)
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods
por: Pennati, Luca, et al.
Publicado: (2026)
por: Pennati, Luca, et al.
Publicado: (2026)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
por: Tramm, John, et al.
Publicado: (2024)
por: Tramm, John, et al.
Publicado: (2024)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
por: Li, Zhouzi, et al.
Publicado: (2026)
por: Li, Zhouzi, et al.
Publicado: (2026)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
por: Hu, Tiancheng, et al.
Publicado: (2026)
por: Hu, Tiancheng, et al.
Publicado: (2026)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
por: Bellavita, Julian, et al.
Publicado: (2026)
por: Bellavita, Julian, et al.
Publicado: (2026)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
por: Peng, Haosong, et al.
Publicado: (2024)
por: Peng, Haosong, et al.
Publicado: (2024)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
por: Costanzo, Manuel, et al.
Publicado: (2024)
por: Costanzo, Manuel, et al.
Publicado: (2024)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
por: Wu, Shixun, et al.
Publicado: (2024)
por: Wu, Shixun, et al.
Publicado: (2024)
A Parallel and Highly-Portable HPC Poisson Solver: Preconditioned Bi-CGSTAB with alpaka
por: Pennati, Luca, et al.
Publicado: (2025)
por: Pennati, Luca, et al.
Publicado: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025)
por: Bhosale, Aditya, et al.
Publicado: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
por: Almasri, Mohammad, et al.
Publicado: (2022)
por: Almasri, Mohammad, et al.
Publicado: (2022)
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
por: Medeiros, Daniel, et al.
Publicado: (2024)
por: Medeiros, Daniel, et al.
Publicado: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
por: Medeiros, Daniel, et al.
Publicado: (2024)
por: Medeiros, Daniel, et al.
Publicado: (2024)
Ejemplares similares
-
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
por: Schieffer, Gabin, et al.
Publicado: (2025) -
QPU Micro-Kernels for Stencil Computation
por: Markidis, Stefano, et al.
Publicado: (2025) -
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
por: Williams, Jeremy J., et al.
Publicado: (2023) -
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
por: Schieffer, Gabin, et al.
Publicado: (2024) -
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)