Learning When to Optimize: Verified Optimization Skills from Expert GPU-Kernel Lineages
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shuoming, Yu, Qiuchu, Zhang, Yangyu, Xu, Ruiyuan, Shi, Xiyu, Li, Guangli, Feng, Xiaobing, Cui, Huimin, Zhao, Jiacheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The New Compiler Stack: A Survey on the Synergy of LLMs and Compilers
by: Zhang, Shuoming, et al.
Published: (2026)
by: Zhang, Shuoming, et al.
Published: (2026)
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
LEGO-Compiler: Enhancing Neural Compilation Through Translation Composability
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
by: Mai, Haohui, et al.
Published: (2026)
by: Mai, Haohui, et al.
Published: (2026)
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
by: Sun, Qitong, et al.
Published: (2026)
by: Sun, Qitong, et al.
Published: (2026)
Beyond Pass-by-Pass Optimization: Intent-Driven IR Optimization with Large Language Models
by: Qiu, Lei, et al.
Published: (2026)
by: Qiu, Lei, et al.
Published: (2026)
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
by: Ke, Changxin, et al.
Published: (2025)
by: Ke, Changxin, et al.
Published: (2025)
Autonomous Sparse Mean-CVaR Portfolio Optimization
by: Lin, Yizun, et al.
Published: (2024)
by: Lin, Yizun, et al.
Published: (2024)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
From Sequential to Parallel: Reformulating Dynamic Programming as GPU Kernels for Large-Scale Stochastic Combinatorial Optimization
by: Zhao, Jingyi, et al.
Published: (2026)
by: Zhao, Jingyi, et al.
Published: (2026)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026)
by: Younesian, Sharareh, et al.
Published: (2026)
Incentivizing LLMs to Self-Verify Their Answers
by: Zhang, Fuxiang, et al.
Published: (2025)
by: Zhang, Fuxiang, et al.
Published: (2025)
Execution-Verified Reinforcement Learning for Optimization Modeling
by: Guan, Runda, et al.
Published: (2026)
by: Guan, Runda, et al.
Published: (2026)
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
by: Won, Jaeyeon, et al.
Published: (2025)
by: Won, Jaeyeon, et al.
Published: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
A Two-Stage GPU Kernel Tuner Combining Semantic Refactoring and Search-Based Optimization
by: Qu, Qiuyi, et al.
Published: (2026)
by: Qu, Qiuyi, et al.
Published: (2026)
Unsupervised Pairwise Learning Optimization Framework for Cross-Corpus EEG-Based Emotion Recognition Based on Prototype Representation
by: Li, Guangli, et al.
Published: (2025)
by: Li, Guangli, et al.
Published: (2025)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
by: Spoczynski, Marcin, et al.
Published: (2026)
by: Spoczynski, Marcin, et al.
Published: (2026)
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
by: Huang, Zixuan, et al.
Published: (2026)
by: Huang, Zixuan, et al.
Published: (2026)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
by: Tschand, Arya, et al.
Published: (2025)
by: Tschand, Arya, et al.
Published: (2025)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
by: Khan, Zaid, et al.
Published: (2026)
by: Khan, Zaid, et al.
Published: (2026)
Recent Advances and Future Prospects for Construction Strategies of Flexible Electromagnetic Protection Patches (Adv. Mater. Technol. 5/2025)
by: De Gong, et al.
Published: (2025)
by: De Gong, et al.
Published: (2025)
Recent Advances and Future Prospects for Construction Strategies of Flexible Electromagnetic Protection Patches
by: De Gong, et al.
Published: (2024)
by: De Gong, et al.
Published: (2024)
Convex Optimization for Alignment and Preference Learning on a Single GPU
by: Feng, Miria, et al.
Published: (2026)
by: Feng, Miria, et al.
Published: (2026)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
by: Hu, Haoyu, et al.
Published: (2026)
by: Hu, Haoyu, et al.
Published: (2026)
Joint Optimization of Maintenance and Quality Inspection for Multi‐Stage Manufacturing System Based on Genetic Reinforcement Learning
by: Haibin Wang, et al.
Published: (2025)
by: Haibin Wang, et al.
Published: (2025)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026)
by: Nichols, Daniel, et al.
Published: (2026)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
FastKernels: Benchmarking GPU Kernel Generation in Production
by: Oliaro, Gabriele, et al.
Published: (2026)
by: Oliaro, Gabriele, et al.
Published: (2026)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2023)
by: Yao, Yihang, et al.
Published: (2023)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
by: Chu, Ruifan, et al.
Published: (2025)
by: Chu, Ruifan, et al.
Published: (2025)
FlipFlop: A Static Analysis-based Energy Optimization Framework for GPU Kernels
by: Rajput, Saurabhsingh, et al.
Published: (2026)
by: Rajput, Saurabhsingh, et al.
Published: (2026)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
by: Zhang, Yijia, et al.
Published: (2024)
by: Zhang, Yijia, et al.
Published: (2024)
GPU-Accelerated Optimization-Based Collision Avoidance
by: Wu, Zeming, et al.
Published: (2024)
by: Wu, Zeming, et al.
Published: (2024)
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations
by: Jia, Fucheng, et al.
Published: (2023)
by: Jia, Fucheng, et al.
Published: (2023)
Similar Items
-
The New Compiler Stack: A Survey on the Synergy of LLMs and Compilers
by: Zhang, Shuoming, et al.
Published: (2026) -
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
by: Zhang, Shuoming, et al.
Published: (2025) -
LEGO-Compiler: Enhancing Neural Compilation Through Translation Composability
by: Zhang, Shuoming, et al.
Published: (2025) -
ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
by: Mai, Haohui, et al.
Published: (2026) -
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
by: Sun, Qitong, et al.
Published: (2026)