Kevin: Multi-Turn RL for Generating CUDA Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Baronio, Carlo, Marsella, Pietro, Pan, Ben, Guo, Simon, Alberti, Silas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025)
by: Zheng, Zhongchun, et al.
Published: (2025)
Lookup multivariate Kolmogorov-Arnold Networks
by: Pozdnyakov, Sergey, et al.
Published: (2025)
by: Pozdnyakov, Sergey, et al.
Published: (2025)
Learning Performance-Improving Code Edits
by: Shypula, Alexander, et al.
Published: (2023)
by: Shypula, Alexander, et al.
Published: (2023)
AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
by: Traini, Luca, et al.
Published: (2024)
by: Traini, Luca, et al.
Published: (2024)
Predicting Configuration Performance in Multiple Environments with Sequential Meta-learning
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2023)
by: Rajput, Saurabhsingh, et al.
Published: (2023)
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
by: Taneja, Jubi, et al.
Published: (2024)
by: Taneja, Jubi, et al.
Published: (2024)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
by: Lange, Robert Tjarko, et al.
Published: (2025)
by: Lange, Robert Tjarko, et al.
Published: (2025)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
by: Saba, Tara, et al.
Published: (2026)
by: Saba, Tara, et al.
Published: (2026)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
by: Wen, Zhongzhen, et al.
Published: (2025)
by: Wen, Zhongzhen, et al.
Published: (2025)
KForge: Program Synthesis for Diverse AI Hardware Accelerators
by: Sereda, Taras, et al.
Published: (2025)
by: Sereda, Taras, et al.
Published: (2025)
Performance Prediction for Large Systems via Text-to-Text Regression
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
by: Yu, Nengneng, et al.
Published: (2026)
by: Yu, Nengneng, et al.
Published: (2026)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
by: Chen, Pengzhou, et al.
Published: (2026)
by: Chen, Pengzhou, et al.
Published: (2026)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
by: Chen, Ziru, et al.
Published: (2026)
by: Chen, Ziru, et al.
Published: (2026)
DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
by: Aghababaeyan, Zohreh, et al.
Published: (2023)
by: Aghababaeyan, Zohreh, et al.
Published: (2023)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
Prompting for Performance: Exploring LLMs for Configuring Software
by: Spieker, Helge, et al.
Published: (2025)
by: Spieker, Helge, et al.
Published: (2025)
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025)
by: Liu, Zhuoran
Published: (2025)
Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software
by: Yi, Lirong, et al.
Published: (2025)
by: Yi, Lirong, et al.
Published: (2025)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis
by: Kumar, Punit, et al.
Published: (2025)
by: Kumar, Punit, et al.
Published: (2025)
Root Cause Localization for Microservice Systems in Cloud-edge Collaborative Environments
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Predicting Software Performance with Divide-and-Learn
by: Gong, Jingzhi, et al.
Published: (2023)
by: Gong, Jingzhi, et al.
Published: (2023)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
by: Rosas, Miguel Romero, et al.
Published: (2024)
by: Rosas, Miguel Romero, et al.
Published: (2024)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
by: Krupp, Lars, et al.
Published: (2026)
by: Krupp, Lars, et al.
Published: (2026)
Automatic Generation of High-Performance RL Environments
by: Karten, Seth, et al.
Published: (2026)
by: Karten, Seth, et al.
Published: (2026)
Prefetching Cache Optimization Using Graph Neural Networks: A Modular Framework and Conceptual Analysis
by: Qowy, F. I.
Published: (2025)
by: Qowy, F. I.
Published: (2025)
Risk-Aware Batch Testing for Performance Regression Detection
by: Sayedsalehi, Ali, et al.
Published: (2026)
by: Sayedsalehi, Ali, et al.
Published: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
by: Wen, Zhongzhen, et al.
Published: (2026)
by: Wen, Zhongzhen, et al.
Published: (2026)
Assuring the Safety of Reinforcement Learning Components: AMLAS-RL
by: Imrie, Calum Corrie, et al.
Published: (2025)
by: Imrie, Calum Corrie, et al.
Published: (2025)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
by: Bhattacharjee, Arijit, et al.
Published: (2026)
by: Bhattacharjee, Arijit, et al.
Published: (2026)
Similar Items
-
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025) -
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025) -
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025) -
Lookup multivariate Kolmogorov-Arnold Networks
by: Pozdnyakov, Sergey, et al.
Published: (2025) -
Learning Performance-Improving Code Edits
by: Shypula, Alexander, et al.
Published: (2023)