KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ran, Dezhi, Xie, Shuxiao, Ji, Mingfang, Liu, Anmin, Wu, Mengzhou, Cao, Yuan, Guo, Yuzhe, Yu, Hao, Li, Linyi, Hu, Yitao, Yang, Wei, Xie, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910018347991040
author Ran, Dezhi
Xie, Shuxiao
Ji, Mingfang
Liu, Anmin
Wu, Mengzhou
Cao, Yuan
Guo, Yuzhe
Yu, Hao
Li, Linyi
Hu, Yitao
Yang, Wei
Xie, Tao
author_facet Ran, Dezhi
Xie, Shuxiao
Ji, Mingfang
Liu, Anmin
Wu, Mengzhou
Cao, Yuan
Guo, Yuzhe
Yu, Hao
Li, Linyi
Hu, Yitao
Yang, Wei
Xie, Tao
contents High-performance GPU kernels are critical for efficient LLM serving, yet their optimization remains a bottleneck requiring deep system expertise. While code LLMs show promise in generating functionally correct code, kernel optimization is intrinsically a search problem over a vast optimization space. The fundamental mismatch prevents existing LLM agents from efficiently exploring the optimization space for diverse hardware and compute patterns. To bridge the gap, we present KernelBand, a framework that formulates kernel optimization as a Multi-Armed Bandit (MAB) problem, explicitly balancing exploration and exploitation to unlock the potential of code LLMs. To navigate the infinite arm space of optimization strategies applied to candidate kernels, we design two key mechanisms: a hardware-aware pruning strategy via profiling bounds and a trace-driven clustering algorithm that leverages Lipschitz continuity. Theoretically, we prove that KernelBand reduces the regret bound to depend on the compact covering number of runtime clusters, ensuring sample-efficient discovery of high-performance kernels. Extensive experiments on TritonBench-G with three GPU architectures and four code LLMs show that KernelBand consistently and substantially outperforms state-of-the-art methods with over 33% average improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18868
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
Ran, Dezhi
Xie, Shuxiao
Ji, Mingfang
Liu, Anmin
Wu, Mengzhou
Cao, Yuan
Guo, Yuzhe
Yu, Hao
Li, Linyi
Hu, Yitao
Yang, Wei
Xie, Tao
Machine Learning
Artificial Intelligence
High-performance GPU kernels are critical for efficient LLM serving, yet their optimization remains a bottleneck requiring deep system expertise. While code LLMs show promise in generating functionally correct code, kernel optimization is intrinsically a search problem over a vast optimization space. The fundamental mismatch prevents existing LLM agents from efficiently exploring the optimization space for diverse hardware and compute patterns. To bridge the gap, we present KernelBand, a framework that formulates kernel optimization as a Multi-Armed Bandit (MAB) problem, explicitly balancing exploration and exploitation to unlock the potential of code LLMs. To navigate the infinite arm space of optimization strategies applied to candidate kernels, we design two key mechanisms: a hardware-aware pruning strategy via profiling bounds and a trace-driven clustering algorithm that leverages Lipschitz continuity. Theoretically, we prove that KernelBand reduces the regret bound to depend on the compact covering number of runtime clusters, ensuring sample-efficient discovery of high-performance kernels. Extensive experiments on TritonBench-G with three GPU architectures and four code LLMs show that KernelBand consistently and substantially outperforms state-of-the-art methods with over 33% average improvement.
title KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.18868