CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Shiyang, Zhang, Zijian, Sun, Guangyan, Luo, Yuebo, Chen, Winson, Wang, Yanzhi, Hong, Mingyi, Ding, Caiwen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
por: Zhang, Zijian, et al.
Publicado: (2025)
por: Zhang, Zijian, et al.
Publicado: (2025)
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
por: Li, Shiyang, et al.
Publicado: (2026)
por: Li, Shiyang, et al.
Publicado: (2026)
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
por: Li, Shiyang, et al.
Publicado: (2026)
por: Li, Shiyang, et al.
Publicado: (2026)
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
por: Luo, Yuebo, et al.
Publicado: (2026)
por: Luo, Yuebo, et al.
Publicado: (2026)
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
por: Han, Yuxuan, et al.
Publicado: (2026)
por: Han, Yuxuan, et al.
Publicado: (2026)
DR-CircuitGNN: Training Acceleration of Heterogeneous Circuit Graph Neural Network on GPUs
por: Luo, Yuebo, et al.
Publicado: (2025)
por: Luo, Yuebo, et al.
Publicado: (2025)
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
por: Zhu, Jiace, et al.
Publicado: (2026)
por: Zhu, Jiace, et al.
Publicado: (2026)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
por: Li, Zhengang, et al.
Publicado: (2024)
por: Li, Zhengang, et al.
Publicado: (2024)
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
por: Chen, Wentao, et al.
Publicado: (2025)
por: Chen, Wentao, et al.
Publicado: (2025)
TROJAN-GUARD: Hardware Trojans Detection Using GNN in RTL Designs
por: Thorat, Kiran, et al.
Publicado: (2025)
por: Thorat, Kiran, et al.
Publicado: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
por: Lange, Robert Tjarko, et al.
Publicado: (2025)
por: Lange, Robert Tjarko, et al.
Publicado: (2025)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
por: Li, Zelin, et al.
Publicado: (2026)
por: Li, Zelin, et al.
Publicado: (2026)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
por: Deng, Kaiyuan, et al.
Publicado: (2026)
por: Deng, Kaiyuan, et al.
Publicado: (2026)
GROOT: Graph Edge Re-growth and Partitioning for the Verification of Large Designs in Logic Synthesis
por: Thorat, Kiran, et al.
Publicado: (2025)
por: Thorat, Kiran, et al.
Publicado: (2025)
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
por: Li, Xiaoya, et al.
Publicado: (2025)
por: Li, Xiaoya, et al.
Publicado: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
Multi-Strategy Improved Snake Optimizer Accelerated CNN-LSTM-Attention-Adaboost for Trajectory Prediction
por: Li, Shiyang
Publicado: (2025)
por: Li, Shiyang
Publicado: (2025)
MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
por: Li, Guangyan, et al.
Publicado: (2025)
por: Li, Guangyan, et al.
Publicado: (2025)
HW-NAS-Bench:Hardware-Aware Neural Architecture Search Benchmark
por: Li, Chaojian, et al.
Publicado: (2021)
por: Li, Chaojian, et al.
Publicado: (2021)
OptiMind: Teaching LLMs to Think Like Optimization Experts
por: Zhang, Xinzhi, et al.
Publicado: (2025)
por: Zhang, Xinzhi, et al.
Publicado: (2025)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
por: Dai, Weinan, et al.
Publicado: (2026)
por: Dai, Weinan, et al.
Publicado: (2026)
Optimizing Sparse Convolution on GPUs with CUDA for 3D Point Cloud Processing in Embedded Systems
por: Luo, Chester, et al.
Publicado: (2024)
por: Luo, Chester, et al.
Publicado: (2024)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
por: Reisizadeh, Hadi, et al.
Publicado: (2025)
por: Reisizadeh, Hadi, et al.
Publicado: (2025)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
por: Zhao, Siyan, et al.
Publicado: (2025)
por: Zhao, Siyan, et al.
Publicado: (2025)
Fortran2CPP: Automating Fortran-to-C++ Translation using LLMs via Multi-Turn Dialogue and Dual-Agent Integration
por: Chen, Le, et al.
Publicado: (2024)
por: Chen, Le, et al.
Publicado: (2024)
DOPPLER: Differentially Private Optimizers with Low-pass Filter for Privacy Noise Reduction
por: Zhang, Xinwei, et al.
Publicado: (2024)
por: Zhang, Xinwei, et al.
Publicado: (2024)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
por: Liu, Mingyi
Publicado: (2026)
por: Liu, Mingyi
Publicado: (2026)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
por: Xie, Xi, et al.
Publicado: (2024)
por: Xie, Xi, et al.
Publicado: (2024)
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
por: Zhang, Yihua, et al.
Publicado: (2024)
por: Zhang, Yihua, et al.
Publicado: (2024)
Attacking All Tasks at Once Using Adversarial Examples in Multi-Task Learning
por: Zhang, Lijun, et al.
Publicado: (2023)
por: Zhang, Lijun, et al.
Publicado: (2023)
Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
Learning to Compare Hardware Designs for High-Level Synthesis
por: Bai, Yunsheng, et al.
Publicado: (2024)
por: Bai, Yunsheng, et al.
Publicado: (2024)
Discretization-independent multifidelity operator learning for partial differential equations
por: Hauck, Jacob, et al.
Publicado: (2025)
por: Hauck, Jacob, et al.
Publicado: (2025)
Hierarchical Mixture of Experts: Generalizable Learning for High-Level Synthesis
por: Li, Weikai, et al.
Publicado: (2024)
por: Li, Weikai, et al.
Publicado: (2024)
Generative modeling assisted simulation of measurement-altered quantum criticality
por: Zhu, Yuchen, et al.
Publicado: (2024)
por: Zhu, Yuchen, et al.
Publicado: (2024)
LLMs Outperform Experts on Challenging Biology Benchmarks
por: Justen, Lennart
Publicado: (2025)
por: Justen, Lennart
Publicado: (2025)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
por: Sukthanker, Rhea Sanjay, et al.
Publicado: (2024)
por: Sukthanker, Rhea Sanjay, et al.
Publicado: (2024)
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
por: Gong, Junfeng, et al.
Publicado: (2025)
por: Gong, Junfeng, et al.
Publicado: (2025)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
por: Zhang, Xinnan, et al.
Publicado: (2025)
por: Zhang, Xinnan, et al.
Publicado: (2025)
Ejemplares similares
-
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
por: Zhang, Zijian, et al.
Publicado: (2025) -
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
por: Li, Shiyang, et al.
Publicado: (2026) -
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
por: Li, Shiyang, et al.
Publicado: (2026) -
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
por: Luo, Yuebo, et al.
Publicado: (2026) -
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
por: Han, Yuxuan, et al.
Publicado: (2026)