STARK: Strategic Team of Agents for Refining Kernels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Juncheng, Yang, Yang, Liu, Tao, Wang, Yang, Qi, Feng, Tarokh, Vahid, Rangadurai, Kaushik, Yang, Shuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918163947454464
author Dong, Juncheng
Yang, Yang
Liu, Tao
Wang, Yang
Qi, Feng
Tarokh, Vahid
Rangadurai, Kaushik
Yang, Shuang
author_facet Dong, Juncheng
Yang, Yang
Liu, Tao
Wang, Yang
Qi, Feng
Tarokh, Vahid
Rangadurai, Kaushik
Yang, Shuang
contents The efficiency of GPU kernels is central to the progress of modern AI, yet optimizing them remains a difficult and labor-intensive task due to complex interactions between memory hierarchies, thread scheduling, and hardware-specific characteristics. While recent advances in large language models (LLMs) provide new opportunities for automated code generation, existing approaches largely treat LLMs as single-shot generators or naive refinement tools, limiting their effectiveness in navigating the irregular kernel optimization landscape. We introduce an LLM agentic framework for GPU kernel optimization that systematically explores the design space through multi-agent collaboration, grounded instruction, dynamic context management, and strategic search. This framework mimics the workflow of expert engineers, enabling LLMs to reason about hardware trade-offs, incorporate profiling feedback, and refine kernels iteratively. We evaluate our approach on KernelBench, a benchmark for LLM-based kernel optimization, and demonstrate substantial improvements over baseline agents: our system produces correct solutions where baselines often fail, and achieves kernels with up to 16x faster runtime performance. These results highlight the potential of agentic LLM frameworks to advance fully automated, scalable GPU kernel optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STARK: Strategic Team of Agents for Refining Kernels
Dong, Juncheng
Yang, Yang
Liu, Tao
Wang, Yang
Qi, Feng
Tarokh, Vahid
Rangadurai, Kaushik
Yang, Shuang
Artificial Intelligence
The efficiency of GPU kernels is central to the progress of modern AI, yet optimizing them remains a difficult and labor-intensive task due to complex interactions between memory hierarchies, thread scheduling, and hardware-specific characteristics. While recent advances in large language models (LLMs) provide new opportunities for automated code generation, existing approaches largely treat LLMs as single-shot generators or naive refinement tools, limiting their effectiveness in navigating the irregular kernel optimization landscape. We introduce an LLM agentic framework for GPU kernel optimization that systematically explores the design space through multi-agent collaboration, grounded instruction, dynamic context management, and strategic search. This framework mimics the workflow of expert engineers, enabling LLMs to reason about hardware trade-offs, incorporate profiling feedback, and refine kernels iteratively. We evaluate our approach on KernelBench, a benchmark for LLM-based kernel optimization, and demonstrate substantial improvements over baseline agents: our system produces correct solutions where baselines often fail, and achieves kernels with up to 16x faster runtime performance. These results highlight the potential of agentic LLM frameworks to advance fully automated, scalable GPU kernel optimization.
title STARK: Strategic Team of Agents for Refining Kernels
topic Artificial Intelligence
url https://arxiv.org/abs/2510.16996