AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Younesian, Sharareh, Ouyang, Wenwen, Rafati, Sina, Rezagholizadeh, Mehdi, Zhou, Sharon, Liu, Ji, Liu, Yue, Yang, Yuchen, Li, Hao, Liu, Ziqiong, Li, Dong, Appia, Vikram, Gu, Zhenyu, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025)
by: Wang, Jianghui, et al.
Published: (2025)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026)
by: Haridas, Akash, et al.
Published: (2026)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
by: Dukler, Yonatan, et al.
Published: (2025)
by: Dukler, Yonatan, et al.
Published: (2025)
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
by: Sun, Qitong, et al.
Published: (2026)
by: Sun, Qitong, et al.
Published: (2026)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
by: An, Zihao, et al.
Published: (2026)
by: An, Zihao, et al.
Published: (2026)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
by: Zhu, Haowei, et al.
Published: (2026)
by: Zhu, Haowei, et al.
Published: (2026)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026)
by: Li, Zekai, et al.
Published: (2026)
DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation
by: jiao, Jiajun, et al.
Published: (2026)
by: jiao, Jiajun, et al.
Published: (2026)
FastKernels: Benchmarking GPU Kernel Generation in Production
by: Oliaro, Gabriele, et al.
Published: (2026)
by: Oliaro, Gabriele, et al.
Published: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
by: Ray, Pretam, et al.
Published: (2026)
by: Ray, Pretam, et al.
Published: (2026)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
by: Wen, Zhongzhen, et al.
Published: (2025)
by: Wen, Zhongzhen, et al.
Published: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Agent Laboratory: Using LLM Agents as Research Assistants
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
by: Ghaffari, Alireza, et al.
Published: (2024)
by: Ghaffari, Alireza, et al.
Published: (2024)
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
by: Ghaffari, Alireza, et al.
Published: (2025)
by: Ghaffari, Alireza, et al.
Published: (2025)
Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
by: Mao, Yuren, et al.
Published: (2025)
by: Mao, Yuren, et al.
Published: (2025)
STARK: Strategic Team of Agents for Refining Kernels
by: Dong, Juncheng, et al.
Published: (2025)
by: Dong, Juncheng, et al.
Published: (2025)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
by: Ji, Haonian, et al.
Published: (2026)
by: Ji, Haonian, et al.
Published: (2026)
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
by: Zhao, Yifan, et al.
Published: (2026)
by: Zhao, Yifan, et al.
Published: (2026)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
by: Singh, Shivam, et al.
Published: (2026)
by: Singh, Shivam, et al.
Published: (2026)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
by: Mishra, Prakamya, et al.
Published: (2025)
by: Mishra, Prakamya, et al.
Published: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
by: Liu, Zefang, et al.
Published: (2025)
by: Liu, Zefang, et al.
Published: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
Learning from Online Videos at Inference Time for Computer-Use Agents
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
NetArena: Dynamic Benchmarks for AI Agents in Network Automation
by: Zhou, Yajie, et al.
Published: (2025)
by: Zhou, Yajie, et al.
Published: (2025)
AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
by: Du, Jinye, et al.
Published: (2025)
by: Du, Jinye, et al.
Published: (2025)
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025)
by: Dubey, Kshitij, et al.
Published: (2025)
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
by: Lin, Edward, et al.
Published: (2026)
by: Lin, Edward, et al.
Published: (2026)
Dual-Kernel Graph Community Contrastive Learning
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Improving Efficiency of GPU Kernel Optimization Agents using a Domain-Specific Language and Speed-of-Light Guidance
by: Hari, Siva Kumar Sastry, et al.
Published: (2026)
by: Hari, Siva Kumar Sastry, et al.
Published: (2026)
Similar Items
-
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025) -
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025) -
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025) -
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026) -
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
by: Dukler, Yonatan, et al.
Published: (2025)