FastKernels: Benchmarking GPU Kernel Generation in Production
Fuente:
arXiv
Saved in:
| Main Authors: | Oliaro, Gabriele, Fu, Yichao, Jiang, May, Lu, Owen, Wang, Junli, Jia, Zhihao, Zhang, Hao, Rajbhandari, Samyam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
by: Qiao, Aurick, et al.
Published: (2024)
by: Qiao, Aurick, et al.
Published: (2024)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026)
by: Younesian, Sharareh, et al.
Published: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
by: Zhang, Zhengxin, et al.
Published: (2024)
by: Zhang, Zhengxin, et al.
Published: (2024)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025)
by: Wang, Jianghui, et al.
Published: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
by: Choi, Hyeong Kyu, et al.
Published: (2025)
by: Choi, Hyeong Kyu, et al.
Published: (2025)
Omniwise: Predicting GPU Kernels Performance with LLMs
by: Wang, Zixian, et al.
Published: (2025)
by: Wang, Zixian, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
by: Sun, Qitong, et al.
Published: (2026)
by: Sun, Qitong, et al.
Published: (2026)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Liger Kernel: Efficient Triton Kernels for LLM Training
by: Hsu, Pin-Lun, et al.
Published: (2024)
by: Hsu, Pin-Lun, et al.
Published: (2024)
AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
by: Du, Weihua, et al.
Published: (2026)
by: Du, Weihua, et al.
Published: (2026)
Model2Kernel: Model-Aware Symbolic Execution For Safe CUDA Kernels
by: He, Mengting, et al.
Published: (2026)
by: He, Mengting, et al.
Published: (2026)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Efficient Estimation of Kernel Surrogate Models for Task Attribution
by: Zhang, Zhenshuo, et al.
Published: (2026)
by: Zhang, Zhenshuo, et al.
Published: (2026)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
ThunderKittens: Simple, Fast, and Adorable AI Kernels
by: Spector, Benjamin F., et al.
Published: (2024)
by: Spector, Benjamin F., et al.
Published: (2024)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
by: Chen, Weizhe, et al.
Published: (2026)
by: Chen, Weizhe, et al.
Published: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
by: Jung, Hee-Jun, et al.
Published: (2022)
by: Jung, Hee-Jun, et al.
Published: (2022)
Fine-Tuning GPT-5 for GPU Kernel Generation
by: Tehrani, Ali, et al.
Published: (2026)
by: Tehrani, Ali, et al.
Published: (2026)
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
by: Lin, Edward, et al.
Published: (2026)
by: Lin, Edward, et al.
Published: (2026)
Understanding Emergent In-Context Learning from a Kernel Regression Perspective
by: Han, Chi, et al.
Published: (2023)
by: Han, Chi, et al.
Published: (2023)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
by: Khan, Zaid, et al.
Published: (2026)
by: Khan, Zaid, et al.
Published: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
Prism: Symbolic Superoptimization of Tensor Programs
by: Wu, Mengdi, et al.
Published: (2026)
by: Wu, Mengdi, et al.
Published: (2026)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
by: Cao, Xinzi, et al.
Published: (2026)
by: Cao, Xinzi, et al.
Published: (2026)
A Quantum Inspired Variational Kernel and Explainable AI Framework for Cross Region Solar and Wind Energy Forecasting
by: Manjunath, Pavan, et al.
Published: (2026)
by: Manjunath, Pavan, et al.
Published: (2026)
Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis
by: Zheng, Yujie, et al.
Published: (2026)
by: Zheng, Yujie, et al.
Published: (2026)
Fast Graph Condensation with Structure-based Neural Tangent Kernel
by: Wang, Lin, et al.
Published: (2023)
by: Wang, Lin, et al.
Published: (2023)
Kernel Banzhaf: A Fast and Robust Estimator for Banzhaf Values
by: Liu, Yurong, et al.
Published: (2024)
by: Liu, Yurong, et al.
Published: (2024)
Similar Items
-
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
by: Qiao, Aurick, et al.
Published: (2024) -
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025) -
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026) -
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024) -
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
by: Pan, Rui, et al.
Published: (2025)