MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Zhongzhen, Zhang, Yinghui, Li, Zhong, Liu, Zhongxin, Xie, Linna, Zhang, Tian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
by: Wen, Zhongzhen, et al.
Published: (2026)
by: Wen, Zhongzhen, et al.
Published: (2026)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
by: Huang, Zixuan, et al.
Published: (2026)
by: Huang, Zixuan, et al.
Published: (2026)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
by: Singh, Samaresh Kumar, et al.
Published: (2026)
by: Singh, Samaresh Kumar, et al.
Published: (2026)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
by: Saba, Tara, et al.
Published: (2026)
by: Saba, Tara, et al.
Published: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective
by: Huang, Mingyu, et al.
Published: (2024)
by: Huang, Mingyu, et al.
Published: (2024)
Should I Run My Cloud Benchmark on Black Friday?
by: Henning, Sören, et al.
Published: (2025)
by: Henning, Sören, et al.
Published: (2025)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
by: Henning, Sören, et al.
Published: (2025)
by: Henning, Sören, et al.
Published: (2025)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
by: Vogel, Adriano, et al.
Published: (2024)
by: Vogel, Adriano, et al.
Published: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
by: Ataie, Ehsan, et al.
Published: (2026)
by: Ataie, Ehsan, et al.
Published: (2026)
Where Should I Deploy My Contracts? A Practical Experience Report
by: Lazăr, Cătălina, et al.
Published: (2025)
by: Lazăr, Cătălina, et al.
Published: (2025)
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)
by: Shipley, Riley, et al.
Published: (2024)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
by: Siavashi, Mohammad, et al.
Published: (2026)
by: Siavashi, Mohammad, et al.
Published: (2026)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
by: Lin, Zhongyi, et al.
Published: (2024)
by: Lin, Zhongyi, et al.
Published: (2024)
MARCO: Multi-Agent Code Optimization with Real-Time Knowledge Integration for High-Performance Computing
by: Rahman, Asif, et al.
Published: (2025)
by: Rahman, Asif, et al.
Published: (2025)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
by: TehraniJamsaz, Ali, et al.
Published: (2024)
by: TehraniJamsaz, Ali, et al.
Published: (2024)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
by: Ruiz, Christian Guzman, et al.
Published: (2024)
by: Ruiz, Christian Guzman, et al.
Published: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
Emergence-as-Code for Self-Governing Reliable Systems
by: Krasnovsky, Anatoly A.
Published: (2026)
by: Krasnovsky, Anatoly A.
Published: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
by: Malhotra, Shubham, et al.
Published: (2025)
by: Malhotra, Shubham, et al.
Published: (2025)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
by: Teranishi, Keita, et al.
Published: (2025)
by: Teranishi, Keita, et al.
Published: (2025)
Adaptive Protein Design Protocols and Middleware
by: Alsaadi, Aymen, et al.
Published: (2025)
by: Alsaadi, Aymen, et al.
Published: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
Root Cause Analysis for Microservice Systems via Cascaded Conditional Learning with Hypergraphs
by: Xie, Shuaiyu, et al.
Published: (2025)
by: Xie, Shuaiyu, et al.
Published: (2025)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
by: Chu, Ruifan, et al.
Published: (2025)
by: Chu, Ruifan, et al.
Published: (2025)
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach
by: Picatto, Hernan, et al.
Published: (2024)
by: Picatto, Hernan, et al.
Published: (2024)
OMPGPT: A Generative Pre-trained Transformer Model for OpenMP
by: Chen, Le, et al.
Published: (2024)
by: Chen, Le, et al.
Published: (2024)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026)
by: Nichols, Daniel, et al.
Published: (2026)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
by: Yang, Peiming, et al.
Published: (2025)
by: Yang, Peiming, et al.
Published: (2025)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
by: Daghero, Francesco, et al.
Published: (2025)
by: Daghero, Francesco, et al.
Published: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
ShuffleBench: A Benchmark for Large-Scale Data Shuffling Operations with Distributed Stream Processing Frameworks
by: Henning, Sören, et al.
Published: (2024)
by: Henning, Sören, et al.
Published: (2024)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
by: Bhattacharjee, Arijit, et al.
Published: (2026)
by: Bhattacharjee, Arijit, et al.
Published: (2026)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
by: Dutta, Akash, et al.
Published: (2024)
by: Dutta, Akash, et al.
Published: (2024)
Similar Items
-
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
by: Wen, Zhongzhen, et al.
Published: (2026) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025) -
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
by: Huang, Zixuan, et al.
Published: (2026) -
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
by: Singh, Samaresh Kumar, et al.
Published: (2026) -
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
by: Saba, Tara, et al.
Published: (2026)