The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
Fuente:
arXiv
Saved in:
| Main Authors: | Amoros, Oscar, Andaluz, Albert, Nunez, Johnny, Pena, Antonio J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
by: Neff, Reece, et al.
Published: (2023)
by: Neff, Reece, et al.
Published: (2023)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
by: Jangda, Abhinav, et al.
Published: (2023)
by: Jangda, Abhinav, et al.
Published: (2023)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
by: Yu, Mingkun, et al.
Published: (2025)
by: Yu, Mingkun, et al.
Published: (2025)
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
by: Zhu, Xinguo, et al.
Published: (2025)
by: Zhu, Xinguo, et al.
Published: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
GPU Sharing with Triples Mode
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
by: Yang, Zhuoping, et al.
Published: (2025)
by: Yang, Zhuoping, et al.
Published: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
Secure API-Driven Research Automation to Accelerate Scientific Discovery
by: Skluzacek, Tyler J., et al.
Published: (2025)
by: Skluzacek, Tyler J., et al.
Published: (2025)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
by: Maureira, Jose, et al.
Published: (2026)
by: Maureira, Jose, et al.
Published: (2026)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
by: Phung, Thanh Son, et al.
Published: (2025)
by: Phung, Thanh Son, et al.
Published: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
DMRlib: Easy-coding and Efficient Resource Management for Job Malleability
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
GPU Programming for AI Workflow Development on AWS SageMaker: An Instructional Approach
by: Srinivasan, Sriram, et al.
Published: (2025)
by: Srinivasan, Sriram, et al.
Published: (2025)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
by: Yu, Zhongming, et al.
Published: (2024)
by: Yu, Zhongming, et al.
Published: (2024)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
by: Park, Seongyeon, et al.
Published: (2024)
by: Park, Seongyeon, et al.
Published: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
by: Doucet, Zachary, et al.
Published: (2025)
by: Doucet, Zachary, et al.
Published: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
by: Xia, Bingzheng, et al.
Published: (2025)
by: Xia, Bingzheng, et al.
Published: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
by: Ranjan, Desh, et al.
Published: (2024)
by: Ranjan, Desh, et al.
Published: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
by: Soos, Dominik, et al.
Published: (2026)
by: Soos, Dominik, et al.
Published: (2026)
Cultivating Multidisciplinary AI Workforce Development on iTiger GPU Cluster: Practices and Challenges
by: Sharif, Mayira, et al.
Published: (2025)
by: Sharif, Mayira, et al.
Published: (2025)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026)
by: Shan, Baodi, et al.
Published: (2026)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
by: Park, Seongyeon, et al.
Published: (2025)
by: Park, Seongyeon, et al.
Published: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
by: Liu, Yunzhao, et al.
Published: (2025)
by: Liu, Yunzhao, et al.
Published: (2025)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
by: Algis, David, et al.
Published: (2024)
by: Algis, David, et al.
Published: (2024)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
by: Latt, Jonas, et al.
Published: (2025)
by: Latt, Jonas, et al.
Published: (2025)
WarpSpeed: A High-Performance Library for Concurrent GPU Hash Tables
by: McCoy, Hunter, et al.
Published: (2025)
by: McCoy, Hunter, et al.
Published: (2025)
Similar Items
-
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023) -
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
by: Neff, Reece, et al.
Published: (2023) -
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025) -
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025) -
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
by: Jangda, Abhinav, et al.
Published: (2023)