KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wiedemann, Nina, Leboutet, Quentin, Paulitsch, Michael, Wofk, Diana, Ummenhofer, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
by: Sul, Stuart H., et al.
Published: (2025)
by: Sul, Stuart H., et al.
Published: (2025)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
by: Wu, Wenqing
Published: (2023)
by: Wu, Wenqing
Published: (2023)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
by: Liu, Xueshen, et al.
Published: (2026)
by: Liu, Xueshen, et al.
Published: (2026)
Fine-Tuning GPT-5 for GPU Kernel Generation
by: Tehrani, Ali, et al.
Published: (2026)
by: Tehrani, Ali, et al.
Published: (2026)
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
GNNBENCH: Fair and Productive Benchmarking for Single-GPU GNN System
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
by: Agrawal, Amey, et al.
Published: (2026)
by: Agrawal, Amey, et al.
Published: (2026)
Keras Sig: Efficient Path Signature Computation on GPU in Keras 3
by: Genet, Rémi, et al.
Published: (2025)
by: Genet, Rémi, et al.
Published: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
by: Shi, Xiaoxiang, et al.
Published: (2025)
by: Shi, Xiaoxiang, et al.
Published: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
by: Masood, Amna, et al.
Published: (2026)
by: Masood, Amna, et al.
Published: (2026)
FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing Maps
by: Xu, Tony, et al.
Published: (2026)
by: Xu, Tony, et al.
Published: (2026)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
by: Mukherjee, Soutrik, et al.
Published: (2026)
by: Mukherjee, Soutrik, et al.
Published: (2026)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
TASP: Topology-aware Sequence Parallelism
by: Wang, Yida, et al.
Published: (2025)
by: Wang, Yida, et al.
Published: (2025)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
by: Shu, Zhihao, et al.
Published: (2026)
by: Shu, Zhihao, et al.
Published: (2026)
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025)
by: Arfeen, Daiyaan, et al.
Published: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
by: Trifan, Octavian Alexandru, et al.
Published: (2025)
by: Trifan, Octavian Alexandru, et al.
Published: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
by: Go, Seokjin, et al.
Published: (2026)
by: Go, Seokjin, et al.
Published: (2026)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
by: Zhao, Lingxiao, et al.
Published: (2025)
by: Zhao, Lingxiao, et al.
Published: (2025)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
by: Yu, Zhongming, et al.
Published: (2024)
by: Yu, Zhongming, et al.
Published: (2024)
Similar Items
-
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
by: Sul, Stuart H., et al.
Published: (2025) -
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
by: Wu, Wenqing
Published: (2023) -
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
by: Liu, Xueshen, et al.
Published: (2026) -
Fine-Tuning GPT-5 for GPU Kernel Generation
by: Tehrani, Ali, et al.
Published: (2026) -
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024)