Stream-K++: Adaptive GPU GEMM Kernel Scheduling and Selection using Bloom Filters
Fuente:
arXiv
Saved in:
| Main Authors: | Sadasivan, Harisankar, Ozturk, Muhammed Emin, Osama, Muhammad, Millette, Chris, Rai, Astha, Podkorytov, Maksim, Afaganis, John, Huang, Carlus, Zhang, Jing, Liu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Bloom Filters for Modern GPU Architectures
by: Jünger, Daniel, et al.
Published: (2025)
by: Jünger, Daniel, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
RadiK: Scalable and Optimized GPU-Parallel Radix Top-K Selection
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks Training
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
QoS-aware Scheduling of Periodic Real-time Task Graphs on Heterogeneous Pre-occupied MECs
by: Shankar, Ashutosh, et al.
Published: (2025)
by: Shankar, Ashutosh, et al.
Published: (2025)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
GPUnion: Autonomous GPU Sharing on Campus
by: Li, Yufang, et al.
Published: (2025)
by: Li, Yufang, et al.
Published: (2025)
Fast Algorithms for Scheduling Many-body Correlation Functions on Accelerators
by: Selvitopi, Oguz, et al.
Published: (2025)
by: Selvitopi, Oguz, et al.
Published: (2025)
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
by: Hu, Zhiyi, et al.
Published: (2025)
by: Hu, Zhiyi, et al.
Published: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
Gathering Semi-Synchronously Scheduled Two-State Robots
by: Otaka, Kohei, et al.
Published: (2024)
by: Otaka, Kohei, et al.
Published: (2024)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
by: Swann, Ryan, et al.
Published: (2024)
by: Swann, Ryan, et al.
Published: (2024)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
by: Da, Wei, et al.
Published: (2025)
by: Da, Wei, et al.
Published: (2025)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
by: Wu, Wenqing
Published: (2023)
by: Wu, Wenqing
Published: (2023)
Theseus: A Distributed and Scalable GPU-Accelerated Query Processing Platform Optimized for Efficient Data Movement
by: Aramburú, Felipe, et al.
Published: (2025)
by: Aramburú, Felipe, et al.
Published: (2025)
$ν$-LPA: Fast GPU-based Label Propagation Algorithm (LPA) for Community Detection
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Analysis of Design Patterns and Benchmark Practices in Apache Kafka Event-Streaming Systems
by: Mohammad, Muzeeb
Published: (2025)
by: Mohammad, Muzeeb
Published: (2025)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
by: Dutta, Yuvraj, et al.
Published: (2025)
by: Dutta, Yuvraj, et al.
Published: (2025)
Efficient GPU Implementation of Static and Incrementally Expanding DF-P PageRank for Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
CPU vs. GPU for Community Detection: Performance Insights from GVE-Louvain and $ν$-Louvain
by: Sahu, Subhajit
Published: (2025)
by: Sahu, Subhajit
Published: (2025)
Memory Efficient GPU-based Label Propagation Algorithm (LPA) for Community Detection on Large Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU
by: Sao, Piyush, et al.
Published: (2024)
by: Sao, Piyush, et al.
Published: (2024)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
by: Maurya, Avinash, et al.
Published: (2025)
by: Maurya, Avinash, et al.
Published: (2025)
Parallel/Distributed Tabu Search for Scheduling Microprocessor Tasks in Hybrid Flowshop
by: Janiak, Adam, et al.
Published: (2025)
by: Janiak, Adam, et al.
Published: (2025)
A Survey on Transactional Stream Processing
by: Zhang, Shuhao, et al.
Published: (2022)
by: Zhang, Shuhao, et al.
Published: (2022)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters
by: Sultana, Abeda, et al.
Published: (2025)
by: Sultana, Abeda, et al.
Published: (2025)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
LP-GEMM: Integrating Layout Propagation into GEMM Operations
by: Carneiro, César Guedes, et al.
Published: (2026)
by: Carneiro, César Guedes, et al.
Published: (2026)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)
by: Pan, Ding, et al.
Published: (2026)
Streaming REST APIs for Large Financial Transaction Exports from Relational Databases
by: Kandiraju, Abhiram
Published: (2026)
by: Kandiraju, Abhiram
Published: (2026)
Similar Items
-
Optimizing Bloom Filters for Modern GPU Architectures
by: Jünger, Daniel, et al.
Published: (2025) -
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025) -
RadiK: Scalable and Optimized GPU-Parallel Radix Top-K Selection
by: Li, Yifei, et al.
Published: (2025) -
MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks Training
by: Peng, Hongwu, et al.
Published: (2023) -
QoS-aware Scheduling of Periodic Real-time Task Graphs on Heterogeneous Pre-occupied MECs
by: Shankar, Ashutosh, et al.
Published: (2025)