Kevin: Multi-Turn RL for Generating CUDA Kernels

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Baronio, Carlo, Marsella, Pietro, Pan, Ben, Guo, Simon, Alberti, Silas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912484981473280
author Baronio, Carlo
Marsella, Pietro
Pan, Ben
Guo, Simon
Alberti, Silas
author_facet Baronio, Carlo
Marsella, Pietro
Pan, Ben
Guo, Simon
Alberti, Silas
contents Writing GPU kernels is a challenging task and critical for AI systems' efficiency. It is also highly iterative: domain experts write code and improve performance through execution feedback. Moreover, it presents verifiable rewards like correctness and speedup, making it a natural environment to apply Reinforcement Learning (RL). To explicitly incorporate the iterative nature of this process into training, we develop a flexible multi-turn RL recipe that addresses unique challenges encountered in real-world settings, such as learning from long trajectories and effective reward attribution across turns. We present Kevin - K(ernel D)evin, the first model trained with multi-turn RL for CUDA kernel generation and optimization. In our evaluation setup, Kevin shows significant gains over its base model (QwQ-32B), improving correctness of generated kernels (in pure CUDA) from 56% to 82% and mean speedup from 0.53x to 1.10x of baseline (PyTorch Eager), and surpassing frontier models like o4-mini (0.78x). Finally, we study its behavior across test-time scaling axes: we found scaling serial refinement more beneficial than parallel sampling. In particular, when given more refinement turns, Kevin shows a higher rate of improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Kevin: Multi-Turn RL for Generating CUDA Kernels
Baronio, Carlo
Marsella, Pietro
Pan, Ben
Guo, Simon
Alberti, Silas
Machine Learning
Artificial Intelligence
Performance
Software Engineering
Writing GPU kernels is a challenging task and critical for AI systems' efficiency. It is also highly iterative: domain experts write code and improve performance through execution feedback. Moreover, it presents verifiable rewards like correctness and speedup, making it a natural environment to apply Reinforcement Learning (RL). To explicitly incorporate the iterative nature of this process into training, we develop a flexible multi-turn RL recipe that addresses unique challenges encountered in real-world settings, such as learning from long trajectories and effective reward attribution across turns. We present Kevin - K(ernel D)evin, the first model trained with multi-turn RL for CUDA kernel generation and optimization. In our evaluation setup, Kevin shows significant gains over its base model (QwQ-32B), improving correctness of generated kernels (in pure CUDA) from 56% to 82% and mean speedup from 0.53x to 1.10x of baseline (PyTorch Eager), and surpassing frontier models like o4-mini (0.78x). Finally, we study its behavior across test-time scaling axes: we found scaling serial refinement more beneficial than parallel sampling. In particular, when given more refinement turns, Kevin shows a higher rate of improvement.
title Kevin: Multi-Turn RL for Generating CUDA Kernels
topic Machine Learning
Artificial Intelligence
Performance
Software Engineering
url https://arxiv.org/abs/2507.11948