Tutoring LLM into a Better CUDA Optimizer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brabec, Matyáš, Klepl, Jiří, Töpfer, Michal, Kruliš, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918163918094336
author Brabec, Matyáš
Klepl, Jiří
Töpfer, Michal
Kruliš, Martin
author_facet Brabec, Matyáš
Klepl, Jiří
Töpfer, Michal
Kruliš, Martin
contents Recent leaps in large language models (LLMs) caused a revolution in programming tools (like GitHub Copilot) that can help with code generation, debugging, and even performance optimization. In this paper, we focus on the capabilities of the most recent reasoning models to generate optimized CUDA code for predefined, well-known tasks. Our objective is to determine which types of code optimizations and parallel patterns the LLMs can perform by themselves and whether they can be improved by tutoring (providing more detailed hints and guidelines in the prompt). The generated solutions were evaluated both automatically (for correctness and speedup) and manually (code reviews) to provide a more detailed perspective. We also tried an interactive approach where the LLM can fix its previous mistakes within a session. The results indicate that LLMs are quite skilled coders; however, they require tutoring to reach optimized solutions provided by parallel computing experts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tutoring LLM into a Better CUDA Optimizer
Brabec, Matyáš
Klepl, Jiří
Töpfer, Michal
Kruliš, Martin
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Recent leaps in large language models (LLMs) caused a revolution in programming tools (like GitHub Copilot) that can help with code generation, debugging, and even performance optimization. In this paper, we focus on the capabilities of the most recent reasoning models to generate optimized CUDA code for predefined, well-known tasks. Our objective is to determine which types of code optimizations and parallel patterns the LLMs can perform by themselves and whether they can be improved by tutoring (providing more detailed hints and guidelines in the prompt). The generated solutions were evaluated both automatically (for correctness and speedup) and manually (code reviews) to provide a more detailed perspective. We also tried an interactive approach where the LLM can fix its previous mistakes within a session. The results indicate that LLMs are quite skilled coders; however, they require tutoring to reach optimized solutions provided by parallel computing experts.
title Tutoring LLM into a Better CUDA Optimizer
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2510.16933