SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Edward, Modi, Sahil, Hari, Siva Kumar Sastry, Huang, Qijing, Ye, Zhifan, Qin, Nestor, Zhou, Fengzhe, Zhang, Yuan, Wang, Jingquan, Damani, Sana, Peri, Dheeraj, Xie, Ouye, Kane, Aditya, Maor, Moshe, Behar, Michael, Cao, Triston, Mehta, Rishabh, Singh, Vartika, Mailthody, Vikram Sharma, Chen, Terry, Ye, Zihao, Chen, Hanfeng, Chen, Tianqi, Grover, Vinod, Chen, Wei, Liu, Wei, Chung, Eric, Ceze, Luis, Bringmann, Roger, Zeller, Cyril, Lightstone, Michael, Kozyrakis, Christos, Shi, Humphrey
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917353640427520
author Lin, Edward
Modi, Sahil
Hari, Siva Kumar Sastry
Huang, Qijing
Ye, Zhifan
Qin, Nestor
Zhou, Fengzhe
Zhang, Yuan
Wang, Jingquan
Damani, Sana
Peri, Dheeraj
Xie, Ouye
Kane, Aditya
Maor, Moshe
Behar, Michael
Cao, Triston
Mehta, Rishabh
Singh, Vartika
Mailthody, Vikram Sharma
Chen, Terry
Ye, Zihao
Chen, Hanfeng
Chen, Tianqi
Grover, Vinod
Chen, Wei
Liu, Wei
Chung, Eric
Ceze, Luis
Bringmann, Roger
Zeller, Cyril
Lightstone, Michael
Kozyrakis, Christos
Shi, Humphrey
author_facet Lin, Edward
Modi, Sahil
Hari, Siva Kumar Sastry
Huang, Qijing
Ye, Zhifan
Qin, Nestor
Zhou, Fengzhe
Zhang, Yuan
Wang, Jingquan
Damani, Sana
Peri, Dheeraj
Xie, Ouye
Kane, Aditya
Maor, Moshe
Behar, Michael
Cao, Triston
Mehta, Rishabh
Singh, Vartika
Mailthody, Vikram Sharma
Chen, Terry
Ye, Zihao
Chen, Hanfeng
Chen, Tianqi
Grover, Vinod
Chen, Wei
Liu, Wei
Chung, Eric
Ceze, Luis
Bringmann, Roger
Zeller, Cyril
Lightstone, Michael
Kozyrakis, Christos
Shi, Humphrey
contents As agentic AI systems become increasingly capable of generating and optimizing GPU kernels, progress is constrained by benchmarks that reward speedup over software baselines rather than proximity to hardware-efficient execution. We present SOL-ExecBench, a benchmark of 235 CUDA kernel optimization problems extracted from 124 production and emerging AI models spanning language, diffusion, vision, audio, video, and hybrid architectures, targeting NVIDIA Blackwell GPUs. The benchmark covers forward and backward workloads across BF16, FP8, and NVFP4, including kernels whose best performance is expected to rely on Blackwell-specific capabilities. Unlike prior benchmarks that evaluate kernels primarily relative to software implementations, SOL-ExecBench measures performance against analytically derived Speed-of-Light (SOL) bounds computed by SOLAR, our pipeline for deriving hardware-grounded SOL bounds, yielding a fixed target for hardware-efficient optimization. We report a SOL Score that quantifies how much of the gap between a release-defined scoring baseline and the hardware SOL bound a candidate kernel closes. To support robust evaluation of agentic optimizers, we additionally provide a sandboxed harness with GPU clock locking, L2 cache clearing, isolated subprocess execution, and static analysis based checks against common reward-hacking strategies. SOL-ExecBench reframes GPU kernel benchmarking from beating a mutable software baseline to closing the remaining gap to hardware Speed-of-Light.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19173
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
Lin, Edward
Modi, Sahil
Hari, Siva Kumar Sastry
Huang, Qijing
Ye, Zhifan
Qin, Nestor
Zhou, Fengzhe
Zhang, Yuan
Wang, Jingquan
Damani, Sana
Peri, Dheeraj
Xie, Ouye
Kane, Aditya
Maor, Moshe
Behar, Michael
Cao, Triston
Mehta, Rishabh
Singh, Vartika
Mailthody, Vikram Sharma
Chen, Terry
Ye, Zihao
Chen, Hanfeng
Chen, Tianqi
Grover, Vinod
Chen, Wei
Liu, Wei
Chung, Eric
Ceze, Luis
Bringmann, Roger
Zeller, Cyril
Lightstone, Michael
Kozyrakis, Christos
Shi, Humphrey
Machine Learning
Artificial Intelligence
As agentic AI systems become increasingly capable of generating and optimizing GPU kernels, progress is constrained by benchmarks that reward speedup over software baselines rather than proximity to hardware-efficient execution. We present SOL-ExecBench, a benchmark of 235 CUDA kernel optimization problems extracted from 124 production and emerging AI models spanning language, diffusion, vision, audio, video, and hybrid architectures, targeting NVIDIA Blackwell GPUs. The benchmark covers forward and backward workloads across BF16, FP8, and NVFP4, including kernels whose best performance is expected to rely on Blackwell-specific capabilities. Unlike prior benchmarks that evaluate kernels primarily relative to software implementations, SOL-ExecBench measures performance against analytically derived Speed-of-Light (SOL) bounds computed by SOLAR, our pipeline for deriving hardware-grounded SOL bounds, yielding a fixed target for hardware-efficient optimization. We report a SOL Score that quantifies how much of the gap between a release-defined scoring baseline and the hardware SOL bound a candidate kernel closes. To support robust evaluation of agentic optimizers, we additionally provide a sandboxed harness with GPU clock locking, L2 cache clearing, isolated subprocess execution, and static analysis based checks against common reward-hacking strategies. SOL-ExecBench reframes GPU kernel benchmarking from beating a mutable software baseline to closing the remaining gap to hardware Speed-of-Light.
title SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.19173