Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Yuxuan, Guo, Meng-Hao, Liu, Zhengning, Chen, Wenguang, Hu, Shi-Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
von: Chen, Wentao, et al.
Veröffentlicht: (2025)
von: Chen, Wentao, et al.
Veröffentlicht: (2025)
CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
von: Li, Shiyang, et al.
Veröffentlicht: (2026)
von: Li, Shiyang, et al.
Veröffentlicht: (2026)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
von: Dai, Weinan, et al.
Veröffentlicht: (2026)
von: Dai, Weinan, et al.
Veröffentlicht: (2026)
Kevin: Multi-Turn RL for Generating CUDA Kernels
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
OptiMind: Teaching LLMs to Think Like Optimization Experts
von: Zhang, Xinzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Xinzhi, et al.
Veröffentlicht: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
von: Guo, Ping, et al.
Veröffentlicht: (2025)
von: Guo, Ping, et al.
Veröffentlicht: (2025)
Spatial-Temporal Mixture-of-Graph-Experts for Multi-Type Crime Prediction
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
von: Bai, Haolei, et al.
Veröffentlicht: (2026)
von: Bai, Haolei, et al.
Veröffentlicht: (2026)
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
von: Zhu, Jiace, et al.
Veröffentlicht: (2026)
von: Zhu, Jiace, et al.
Veröffentlicht: (2026)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
von: Dong, Kris Shengjun, et al.
Veröffentlicht: (2026)
von: Dong, Kris Shengjun, et al.
Veröffentlicht: (2026)
TiledAttention: a CUDA Tile SDPA Kernel for PyTorch
von: Khan, Taimur
Veröffentlicht: (2026)
von: Khan, Taimur
Veröffentlicht: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
von: Sun, Qitong, et al.
Veröffentlicht: (2026)
von: Sun, Qitong, et al.
Veröffentlicht: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2026)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2026)
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
Collaborative Expert LLMs Guided Multi-Objective Molecular Optimization
von: Yu, Jiajun, et al.
Veröffentlicht: (2025)
von: Yu, Jiajun, et al.
Veröffentlicht: (2025)
PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
MixTTE: Multi-Level Mixture-of-Experts for Scalable and Adaptive Travel Time Estimation
von: Jiang, Wenzhao, et al.
Veröffentlicht: (2026)
von: Jiang, Wenzhao, et al.
Veröffentlicht: (2026)
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
von: Guo, Jingming, et al.
Veröffentlicht: (2024)
von: Guo, Jingming, et al.
Veröffentlicht: (2024)
KernelBench: Can LLMs Write Efficient GPU Kernels?
von: Ouyang, Anne, et al.
Veröffentlicht: (2025)
von: Ouyang, Anne, et al.
Veröffentlicht: (2025)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Exploring Critical Testing Scenarios for Decision-Making Policies: An LLM Approach
von: Xu, Weichao, et al.
Veröffentlicht: (2024)
von: Xu, Weichao, et al.
Veröffentlicht: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
von: Du, He, et al.
Veröffentlicht: (2026)
von: Du, He, et al.
Veröffentlicht: (2026)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
von: Li, Albus Yizhuo
Veröffentlicht: (2025)
von: Li, Albus Yizhuo
Veröffentlicht: (2025)
MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
von: Zou, Xingze, et al.
Veröffentlicht: (2026)
von: Zou, Xingze, et al.
Veröffentlicht: (2026)
Lightweight Gaussian Process Inference in C++ on Metal and CUDA
von: Fang, Yu-Hsueh
Veröffentlicht: (2026)
von: Fang, Yu-Hsueh
Veröffentlicht: (2026)
AdaKernel: Learning Adaptive Kernel Parameters for Spatiotemporal Graph Neural Networks
von: Zhang, Zhongyue, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongyue, et al.
Veröffentlicht: (2026)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
Tight Clusters Make Specialized Experts
von: Nielsen, Stefan K., et al.
Veröffentlicht: (2025)
von: Nielsen, Stefan K., et al.
Veröffentlicht: (2025)
Towards Automated Kernel Generation in the Era of LLMs
von: Yu, Yang, et al.
Veröffentlicht: (2026)
von: Yu, Yang, et al.
Veröffentlicht: (2026)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
von: Chen, Wentao, et al.
Veröffentlicht: (2025) -
CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
von: Li, Shiyang, et al.
Veröffentlicht: (2026) -
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
von: Dai, Weinan, et al.
Veröffentlicht: (2026) -
Kevin: Multi-Turn RL for Generating CUDA Kernels
von: Baronio, Carlo, et al.
Veröffentlicht: (2025) -
OptiMind: Teaching LLMs to Think Like Optimization Experts
von: Zhang, Xinzhi, et al.
Veröffentlicht: (2025)