FastKernels: Benchmarking GPU Kernel Generation in Production
Fuente:
arXiv
Guardado en:
| Autores principales: | Oliaro, Gabriele, Fu, Yichao, Jiang, May, Lu, Owen, Wang, Junli, Jia, Zhihao, Zhang, Hao, Rajbhandari, Samyam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
por: Qiao, Aurick, et al.
Publicado: (2024)
por: Qiao, Aurick, et al.
Publicado: (2024)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
por: Lee, Jaeseong, et al.
Publicado: (2025)
por: Lee, Jaeseong, et al.
Publicado: (2025)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
por: Younesian, Sharareh, et al.
Publicado: (2026)
por: Younesian, Sharareh, et al.
Publicado: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
por: Pan, Rui, et al.
Publicado: (2025)
por: Pan, Rui, et al.
Publicado: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
por: Liu, Wei, et al.
Publicado: (2026)
por: Liu, Wei, et al.
Publicado: (2026)
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
por: Zhang, Zhengxin, et al.
Publicado: (2024)
por: Zhang, Zhengxin, et al.
Publicado: (2024)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
por: Wang, Jianghui, et al.
Publicado: (2025)
por: Wang, Jianghui, et al.
Publicado: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
por: Yang, Lijie, et al.
Publicado: (2024)
por: Yang, Lijie, et al.
Publicado: (2024)
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
por: Choi, Hyeong Kyu, et al.
Publicado: (2025)
por: Choi, Hyeong Kyu, et al.
Publicado: (2025)
Omniwise: Predicting GPU Kernels Performance with LLMs
por: Wang, Zixian, et al.
Publicado: (2025)
por: Wang, Zixian, et al.
Publicado: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
por: Ouyang, Anne, et al.
Publicado: (2025)
por: Ouyang, Anne, et al.
Publicado: (2025)
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
por: Sun, Qitong, et al.
Publicado: (2026)
por: Sun, Qitong, et al.
Publicado: (2026)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
por: Miao, Xupeng, et al.
Publicado: (2023)
por: Miao, Xupeng, et al.
Publicado: (2023)
Liger Kernel: Efficient Triton Kernels for LLM Training
por: Hsu, Pin-Lun, et al.
Publicado: (2024)
por: Hsu, Pin-Lun, et al.
Publicado: (2024)
AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
por: Du, Weihua, et al.
Publicado: (2026)
por: Du, Weihua, et al.
Publicado: (2026)
Model2Kernel: Model-Aware Symbolic Execution For Safe CUDA Kernels
por: He, Mengting, et al.
Publicado: (2026)
por: He, Mengting, et al.
Publicado: (2026)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
Efficient Estimation of Kernel Surrogate Models for Task Attribution
por: Zhang, Zhenshuo, et al.
Publicado: (2026)
por: Zhang, Zhenshuo, et al.
Publicado: (2026)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
ThunderKittens: Simple, Fast, and Adorable AI Kernels
por: Spector, Benjamin F., et al.
Publicado: (2024)
por: Spector, Benjamin F., et al.
Publicado: (2024)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
por: Das, Amitava, et al.
Publicado: (2025)
por: Das, Amitava, et al.
Publicado: (2025)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
por: Chen, Weizhe, et al.
Publicado: (2026)
por: Chen, Weizhe, et al.
Publicado: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
por: Jung, Hee-Jun, et al.
Publicado: (2022)
por: Jung, Hee-Jun, et al.
Publicado: (2022)
Fine-Tuning GPT-5 for GPU Kernel Generation
por: Tehrani, Ali, et al.
Publicado: (2026)
por: Tehrani, Ali, et al.
Publicado: (2026)
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
por: Lin, Edward, et al.
Publicado: (2026)
por: Lin, Edward, et al.
Publicado: (2026)
Understanding Emergent In-Context Learning from a Kernel Regression Perspective
por: Han, Chi, et al.
Publicado: (2023)
por: Han, Chi, et al.
Publicado: (2023)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
por: Kim, Taesu, et al.
Publicado: (2024)
por: Kim, Taesu, et al.
Publicado: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
por: Wang, Cheng, et al.
Publicado: (2026)
por: Wang, Cheng, et al.
Publicado: (2026)
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
por: Hu, Lanxiang, et al.
Publicado: (2025)
por: Hu, Lanxiang, et al.
Publicado: (2025)
GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
por: Khan, Zaid, et al.
Publicado: (2026)
por: Khan, Zaid, et al.
Publicado: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025)
por: Li, Zikun, et al.
Publicado: (2025)
Prism: Symbolic Superoptimization of Tensor Programs
por: Wu, Mengdi, et al.
Publicado: (2026)
por: Wu, Mengdi, et al.
Publicado: (2026)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
por: Nikitin, Alexander, et al.
Publicado: (2024)
por: Nikitin, Alexander, et al.
Publicado: (2024)
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
por: Cao, Xinzi, et al.
Publicado: (2026)
por: Cao, Xinzi, et al.
Publicado: (2026)
A Quantum Inspired Variational Kernel and Explainable AI Framework for Cross Region Solar and Wind Energy Forecasting
por: Manjunath, Pavan, et al.
Publicado: (2026)
por: Manjunath, Pavan, et al.
Publicado: (2026)
Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis
por: Zheng, Yujie, et al.
Publicado: (2026)
por: Zheng, Yujie, et al.
Publicado: (2026)
Fast Graph Condensation with Structure-based Neural Tangent Kernel
por: Wang, Lin, et al.
Publicado: (2023)
por: Wang, Lin, et al.
Publicado: (2023)
Kernel Banzhaf: A Fast and Robust Estimator for Banzhaf Values
por: Liu, Yurong, et al.
Publicado: (2024)
por: Liu, Yurong, et al.
Publicado: (2024)
Ejemplares similares
-
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
por: Qiao, Aurick, et al.
Publicado: (2024) -
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
por: Lee, Jaeseong, et al.
Publicado: (2025) -
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
por: Younesian, Sharareh, et al.
Publicado: (2026) -
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
por: Oliaro, Gabriele, et al.
Publicado: (2024) -
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
por: Pan, Rui, et al.
Publicado: (2025)