TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Jianling, Li, Shangzhan, Gao, Zhenye, Shi, Qi, Li, Yuxuan, Wang, Zefan, Huang, Jiacheng, Wang, Haojie, Wang, Jianrong, Han, Xu, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917931085987840
author Li, Jianling
Li, Shangzhan
Gao, Zhenye
Shi, Qi
Li, Yuxuan
Wang, Zefan
Huang, Jiacheng
Wang, Haojie
Wang, Jianrong
Han, Xu
Liu, Zhiyuan
Sun, Maosong
author_facet Li, Jianling
Li, Shangzhan
Gao, Zhenye
Shi, Qi
Li, Yuxuan
Wang, Zefan
Huang, Jiacheng
Wang, Haojie
Wang, Jianrong
Han, Xu
Liu, Zhiyuan
Sun, Maosong
contents Triton, a high-level Python-like language designed for building efficient GPU kernels, is widely adopted in deep learning frameworks due to its portability, flexibility, and accessibility. However, programming and parallel optimization still require considerable trial and error from Triton developers. Despite advances in large language models (LLMs) for conventional code generation, these models struggle to generate accurate, performance-optimized Triton code, as they lack awareness of its specifications and the complexities of GPU programming. More critically, there is an urgent need for systematic evaluations tailored to Triton. In this work, we introduce TritonBench, the first comprehensive benchmark for Triton operator generation. TritonBench features two evaluation channels: a curated set of 184 real-world operators from GitHub and a collection of operators aligned with PyTorch interfaces. Unlike conventional code benchmarks prioritizing functional correctness, TritonBench also profiles efficiency performance on widely deployed GPUs aligned with industry applications. Our study reveals that current state-of-the-art code LLMs struggle to generate efficient Triton operators, highlighting a significant gap in high-performance code generation. TritonBench will be available at https://github.com/thunlp/TritonBench.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14752
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
Li, Jianling
Li, Shangzhan
Gao, Zhenye
Shi, Qi
Li, Yuxuan
Wang, Zefan
Huang, Jiacheng
Wang, Haojie
Wang, Jianrong
Han, Xu
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Machine Learning
Triton, a high-level Python-like language designed for building efficient GPU kernels, is widely adopted in deep learning frameworks due to its portability, flexibility, and accessibility. However, programming and parallel optimization still require considerable trial and error from Triton developers. Despite advances in large language models (LLMs) for conventional code generation, these models struggle to generate accurate, performance-optimized Triton code, as they lack awareness of its specifications and the complexities of GPU programming. More critically, there is an urgent need for systematic evaluations tailored to Triton. In this work, we introduce TritonBench, the first comprehensive benchmark for Triton operator generation. TritonBench features two evaluation channels: a curated set of 184 real-world operators from GitHub and a collection of operators aligned with PyTorch interfaces. Unlike conventional code benchmarks prioritizing functional correctness, TritonBench also profiles efficiency performance on widely deployed GPUs aligned with industry applications. Our study reveals that current state-of-the-art code LLMs struggle to generate efficient Triton operators, highlighting a significant gap in high-performance code generation. TritonBench will be available at https://github.com/thunlp/TritonBench.
title TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.14752