KEET: Explaining Performance of GPU Kernels Using LLM Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Davis, Joshua H., Rydzy, Klaudiusz, Ramesh, Srinivasan, Nilay, Aadit, Nichols, Daniel, Raj, Swapna, Jain, Nikhil, Bhatele, Abhinav
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911652101750784
author Davis, Joshua H.
Rydzy, Klaudiusz
Ramesh, Srinivasan
Nilay, Aadit
Nichols, Daniel
Raj, Swapna
Jain, Nikhil
Bhatele, Abhinav
author_facet Davis, Joshua H.
Rydzy, Klaudiusz
Ramesh, Srinivasan
Nilay, Aadit
Nichols, Daniel
Raj, Swapna
Jain, Nikhil
Bhatele, Abhinav
contents Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU architecture, kernel developers need to spend significant time analyzing and comparing profiles in the tool's graphical interface to identify and understand kernel performance bottlenecks. Large Language Models (LLMs) have shown promise in understanding complex data and generating natural language explanations. In this paper, we propose the Kernel Execution Explanation Toolkit (KEET), an LLM-based agentic framework for interpreting Nsight Compute profiles to generate useful and data-grounded natural language explanations of performance issues in GPU kernels, and suggestions for optimizations. We evaluate \toolname using several CUDA kernels of varying complexity on NVIDIA H100 GPUs. We find that the generated explanations, when provided as context, improve the quality of LLM code optimization and multiple-choice question answering in downstream tasks. We further demonstrate that the tool can be used to interpret performance data from large sets of profiles to improve the quality of optimization suggestions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04467
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KEET: Explaining Performance of GPU Kernels Using LLM Agents
Davis, Joshua H.
Rydzy, Klaudiusz
Ramesh, Srinivasan
Nilay, Aadit
Nichols, Daniel
Raj, Swapna
Jain, Nikhil
Bhatele, Abhinav
Performance
Distributed, Parallel, and Cluster Computing
Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU architecture, kernel developers need to spend significant time analyzing and comparing profiles in the tool's graphical interface to identify and understand kernel performance bottlenecks. Large Language Models (LLMs) have shown promise in understanding complex data and generating natural language explanations. In this paper, we propose the Kernel Execution Explanation Toolkit (KEET), an LLM-based agentic framework for interpreting Nsight Compute profiles to generate useful and data-grounded natural language explanations of performance issues in GPU kernels, and suggestions for optimizations. We evaluate \toolname using several CUDA kernels of varying complexity on NVIDIA H100 GPUs. We find that the generated explanations, when provided as context, improve the quality of LLM code optimization and multiple-choice question answering in downstream tasks. We further demonstrate that the tool can be used to interpret performance data from large sets of profiles to improve the quality of optimization suggestions.
title KEET: Explaining Performance of GPU Kernels Using LLM Agents
topic Performance
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.04467