SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tschand, Arya, Awad, Muhammad, Swann, Ryan, Ramakrishnan, Kesavan, Ma, Jeffrey, Lowery, Keith, Dasika, Ganesh, Reddi, Vijay Janapa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914010663747584
author Tschand, Arya
Awad, Muhammad
Swann, Ryan
Ramakrishnan, Kesavan
Ma, Jeffrey
Lowery, Keith
Dasika, Ganesh
Reddi, Vijay Janapa
author_facet Tschand, Arya
Awad, Muhammad
Swann, Ryan
Ramakrishnan, Kesavan
Ma, Jeffrey
Lowery, Keith
Dasika, Ganesh
Reddi, Vijay Janapa
contents Large language models (LLMs) have shown progress in GPU kernel performance engineering using inefficient search-based methods that optimize around runtime. Any existing approach lacks a key characteristic that human performance engineers rely on for near-optimal utilization -- hardware-awareness. By leveraging the workload's specific memory access patterns, architecture specifications, filtered profiling logs, and reflections on historical performance, we can make software-level optimizations that are tailored to the underlying hardware. SwizzlePerf automatically generates spatial optimizations for GPU kernels on disaggregated architectures by giving LLMs explicit hardware-awareness. For a GEMM kernel, SwizzlePerf takes less than 5 minutes to generate the same hardware-specific optimal swizzling pattern that took expert performance engineers 2 weeks to find. On a suite of 10 diverse ML and Science kernels, SwizzlePerf can generate swizzling patterns for 9 of the kernels that achieve up to a 2.06x speedup and 70% improvement in L2 hit rate. This work is the first of many steps toward systematically creating hardware-aware LLM performance engineering agents.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20258
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
Tschand, Arya
Awad, Muhammad
Swann, Ryan
Ramakrishnan, Kesavan
Ma, Jeffrey
Lowery, Keith
Dasika, Ganesh
Reddi, Vijay Janapa
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Large language models (LLMs) have shown progress in GPU kernel performance engineering using inefficient search-based methods that optimize around runtime. Any existing approach lacks a key characteristic that human performance engineers rely on for near-optimal utilization -- hardware-awareness. By leveraging the workload's specific memory access patterns, architecture specifications, filtered profiling logs, and reflections on historical performance, we can make software-level optimizations that are tailored to the underlying hardware. SwizzlePerf automatically generates spatial optimizations for GPU kernels on disaggregated architectures by giving LLMs explicit hardware-awareness. For a GEMM kernel, SwizzlePerf takes less than 5 minutes to generate the same hardware-specific optimal swizzling pattern that took expert performance engineers 2 weeks to find. On a suite of 10 diverse ML and Science kernels, SwizzlePerf can generate swizzling patterns for 9 of the kernels that achieve up to a 2.06x speedup and 70% improvement in L2 hit rate. This work is the first of many steps toward systematically creating hardware-aware LLM performance engineering agents.
title SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2508.20258