KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Kris Shengjun, Modi, Sahil, Nikiforov, Dima, Damani, Sana, Lin, Edward, Hari, Siva Kumar Sastry, Kozyrakis, Christos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProofWright: Towards Agentic Formal Verification of CUDA
von: Chatterjee, Bodhisatwa, et al.
Veröffentlicht: (2025)
von: Chatterjee, Bodhisatwa, et al.
Veröffentlicht: (2025)
Improving Efficiency of GPU Kernel Optimization Agents using a Domain-Specific Language and Speed-of-Light Guidance
von: Hari, Siva Kumar Sastry, et al.
Veröffentlicht: (2026)
von: Hari, Siva Kumar Sastry, et al.
Veröffentlicht: (2026)
Hunting CUDA Bugs at Scale with cuFuzz
von: ziad, Mohamed Tarek Ibn, et al.
Veröffentlicht: (2026)
von: ziad, Mohamed Tarek Ibn, et al.
Veröffentlicht: (2026)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
LLM-Aided Compilation for Tensor Accelerators
von: Hong, Charles, et al.
Veröffentlicht: (2024)
von: Hong, Charles, et al.
Veröffentlicht: (2024)
Safety-Critical Scenario Generation Via Reinforcement Learning Based Editing
von: Liu, Haolan, et al.
Veröffentlicht: (2023)
von: Liu, Haolan, et al.
Veröffentlicht: (2023)
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
von: Lin, Edward, et al.
Veröffentlicht: (2026)
von: Lin, Edward, et al.
Veröffentlicht: (2026)
Characterizing and Optimizing Real-Time Optimal Control for Embedded SoCs
von: Dong, Kris Shengjun, et al.
Veröffentlicht: (2024)
von: Dong, Kris Shengjun, et al.
Veröffentlicht: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
von: Chen, Wentao, et al.
Veröffentlicht: (2025)
von: Chen, Wentao, et al.
Veröffentlicht: (2025)
ALBERTA: ALgorithm-Based Error Resilience in Transformer Architectures
von: Liu, Haoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Haoxuan, et al.
Veröffentlicht: (2023)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
von: Dai, Weinan, et al.
Veröffentlicht: (2026)
von: Dai, Weinan, et al.
Veröffentlicht: (2026)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
von: Ziad, Mohamed Tarek Ibn, et al.
Veröffentlicht: (2025)
von: Ziad, Mohamed Tarek Ibn, et al.
Veröffentlicht: (2025)
AutoTask: Task Aware Multi-Faceted Single Model for Multi-Task Ads Relevance
von: Guo, Shouchang, et al.
Veröffentlicht: (2024)
von: Guo, Shouchang, et al.
Veröffentlicht: (2024)
Conversation Kernels: A Flexible Mechanism to Learn Relevant Context for Online Conversation Understanding
von: Agarwal, Vibhor, et al.
Veröffentlicht: (2025)
von: Agarwal, Vibhor, et al.
Veröffentlicht: (2025)
Black Blaster: A Symbiotic Electromagnetic Weapon Based on Palindromic Fractal Equilibrium
von: Fuertes Oliva, Martín
Veröffentlicht: (2025)
von: Fuertes Oliva, Martín
Veröffentlicht: (2025)
Black Blaster: A Symbiotic Electromagnetic Weapon Based on Palindromic Fractal Equilibrium
von: Fuertes Oliva, Martín
Veröffentlicht: (2025)
von: Fuertes Oliva, Martín
Veröffentlicht: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
von: Jiang, Wenqi, et al.
Veröffentlicht: (2026)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2026)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
von: Balaji, Vignesh, et al.
Veröffentlicht: (2025)
von: Balaji, Vignesh, et al.
Veröffentlicht: (2025)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Model2Kernel: Model-Aware Symbolic Execution For Safe CUDA Kernels
von: He, Mengting, et al.
Veröffentlicht: (2026)
von: He, Mengting, et al.
Veröffentlicht: (2026)
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
von: Wu, Meng, et al.
Veröffentlicht: (2024)
von: Wu, Meng, et al.
Veröffentlicht: (2024)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
Self-Distillation Enables Continual Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
TiledAttention: a CUDA Tile SDPA Kernel for PyTorch
von: Khan, Taimur
Veröffentlicht: (2026)
von: Khan, Taimur
Veröffentlicht: (2026)
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
von: Han, Yuxuan, et al.
Veröffentlicht: (2026)
von: Han, Yuxuan, et al.
Veröffentlicht: (2026)
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
von: Bai, Haolei, et al.
Veröffentlicht: (2026)
von: Bai, Haolei, et al.
Veröffentlicht: (2026)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
Certifiable Reachability Learning Using a New Lipschitz Continuous Value Function
von: Li, Jingqi, et al.
Veröffentlicht: (2024)
von: Li, Jingqi, et al.
Veröffentlicht: (2024)
Teaching Cloud Infrastructure and Scalable Application Deployment in an Undergraduate Computer Science Program
von: Saligrama, Aditya, et al.
Veröffentlicht: (2024)
von: Saligrama, Aditya, et al.
Veröffentlicht: (2024)
State Contamination in Memory-Augmented LLM Agents
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
GPU-RANC: A CUDA Accelerated Simulation Framework for Neuromorphic Architectures
von: Hassan, Sahil, et al.
Veröffentlicht: (2024)
von: Hassan, Sahil, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ProofWright: Towards Agentic Formal Verification of CUDA
von: Chatterjee, Bodhisatwa, et al.
Veröffentlicht: (2025) -
Improving Efficiency of GPU Kernel Optimization Agents using a Domain-Specific Language and Speed-of-Light Guidance
von: Hari, Siva Kumar Sastry, et al.
Veröffentlicht: (2026) -
Hunting CUDA Bugs at Scale with cuFuzz
von: ziad, Mohamed Tarek Ibn, et al.
Veröffentlicht: (2026) -
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026) -
LLM-Aided Compilation for Tensor Accelerators
von: Hong, Charles, et al.
Veröffentlicht: (2024)