Improving the Efficiency of OpenCL Kernels through Pipes
Fuente:
arXiv
Saved in:
| Main Authors: | Zarch, Mostafa Eghbali, Becchi, Michela |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
by: Neff, Reece, et al.
Published: (2023)
by: Neff, Reece, et al.
Published: (2023)
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
by: Guidotti, Nicolas Lee
Published: (2025)
by: Guidotti, Nicolas Lee
Published: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing
by: Ferdous, S M, et al.
Published: (2024)
by: Ferdous, S M, et al.
Published: (2024)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
by: Liu, Chongpeng, et al.
Published: (2025)
by: Liu, Chongpeng, et al.
Published: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
by: Ramachandran, Arun, et al.
Published: (2025)
by: Ramachandran, Arun, et al.
Published: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
by: Wang, Hongyu, et al.
Published: (2026)
by: Wang, Hongyu, et al.
Published: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
by: Wang, Weiye, et al.
Published: (2026)
by: Wang, Weiye, et al.
Published: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
by: Lin, Yanying, et al.
Published: (2025)
by: Lin, Yanying, et al.
Published: (2025)
Redox: Improving I/O Efficiency of Model Training Through File Redirection
by: Li, Yuhao, et al.
Published: (2025)
by: Li, Yuhao, et al.
Published: (2025)
Enhancing Energy Efficiency in Scientific Workflows through CFD based PIVAEs
by: Zahir, Ali, et al.
Published: (2026)
by: Zahir, Ali, et al.
Published: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
by: Jiang, Chaoyi, et al.
Published: (2024)
by: Jiang, Chaoyi, et al.
Published: (2024)
Serverless Computing: Architecture, Concepts, and Applications
by: Ghorbian, Mohsen, et al.
Published: (2025)
by: Ghorbian, Mohsen, et al.
Published: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
On Similarity of Computational Kernels in our Codes and Proxies
by: McKinsey, Michael, et al.
Published: (2026)
by: McKinsey, Michael, et al.
Published: (2026)
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
by: Swann, Ryan, et al.
Published: (2024)
by: Swann, Ryan, et al.
Published: (2024)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
by: Jangda, Abhinav, et al.
Published: (2023)
by: Jangda, Abhinav, et al.
Published: (2023)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
by: Zheng, Size, et al.
Published: (2026)
by: Zheng, Size, et al.
Published: (2026)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
by: Amoros, Oscar, et al.
Published: (2025)
by: Amoros, Oscar, et al.
Published: (2025)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
by: Peng, Ivy, et al.
Published: (2024)
by: Peng, Ivy, et al.
Published: (2024)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
by: Fridman, Yehonatan, et al.
Published: (2024)
by: Fridman, Yehonatan, et al.
Published: (2024)
Portability Efficiency Approach for Calculating Performance Portability
by: Marowka, Ami
Published: (2024)
by: Marowka, Ami
Published: (2024)
Similar Items
-
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
by: Neff, Reece, et al.
Published: (2023) -
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
by: Guidotti, Nicolas Lee
Published: (2025) -
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026) -
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024) -
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)