MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wen, Zhongzhen, Zhang, Yinghui, Li, Zhong, Liu, Zhongxin, Xie, Linna, Zhang, Tian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908468088143872
author Wen, Zhongzhen
Zhang, Yinghui
Li, Zhong
Liu, Zhongxin
Xie, Linna
Zhang, Tian
author_facet Wen, Zhongzhen
Zhang, Yinghui
Li, Zhong
Liu, Zhongxin
Xie, Linna
Zhang, Tian
contents The automatic generation of deep learning (DL) kernels using large language models (LLMs) has emerged as a promising approach to reduce the manual effort and hardware-specific expertise required for writing high-performance operator implementations. However, existing benchmarks for evaluating LLMs in this domain suffer from limited hardware support, coarse-grained kernel categorization, and imbalanced task coverage. To address these limitations, we introduce MultiKernelBench, the first comprehensive, multi-platform benchmark for LLM-based DL kernel generation. MultiKernelBench spans 285 tasks across 14 well-defined kernel categories and supports three major hardware platforms: Nvidia GPUs, Huawei NPUs, and Google TPUs. To enable future extensibility, we design a modular backend abstraction layer that decouples platform-specific logic from the core benchmarking infrastructure, allowing easy integration of new hardware platforms. We further propose a simple yet effective category-aware one-shot prompting method that improves generation quality by providing in-category exemplars. Through systematic evaluations of seven state-of-the-art LLMs, we reveal significant variation in task difficulty, poor generalization to platforms with less training exposure, and the effectiveness of targeted prompting strategies. MultiKernelBench is publicly available at https://github.com/wzzll123/MultiKernelBench.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
Wen, Zhongzhen
Zhang, Yinghui
Li, Zhong
Liu, Zhongxin
Xie, Linna
Zhang, Tian
Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
Software Engineering
The automatic generation of deep learning (DL) kernels using large language models (LLMs) has emerged as a promising approach to reduce the manual effort and hardware-specific expertise required for writing high-performance operator implementations. However, existing benchmarks for evaluating LLMs in this domain suffer from limited hardware support, coarse-grained kernel categorization, and imbalanced task coverage. To address these limitations, we introduce MultiKernelBench, the first comprehensive, multi-platform benchmark for LLM-based DL kernel generation. MultiKernelBench spans 285 tasks across 14 well-defined kernel categories and supports three major hardware platforms: Nvidia GPUs, Huawei NPUs, and Google TPUs. To enable future extensibility, we design a modular backend abstraction layer that decouples platform-specific logic from the core benchmarking infrastructure, allowing easy integration of new hardware platforms. We further propose a simple yet effective category-aware one-shot prompting method that improves generation quality by providing in-category exemplars. Through systematic evaluations of seven state-of-the-art LLMs, we reveal significant variation in task difficulty, poor generalization to platforms with less training exposure, and the effectiveness of targeted prompting strategies. MultiKernelBench is publicly available at https://github.com/wzzll123/MultiKernelBench.
title MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
topic Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
Software Engineering
url https://arxiv.org/abs/2507.17773