CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
Fuente:
arXiv
Guardado en:
| Autores principales: | Saba, Tara, Ouyang, Anne, Si, Xujie, Long, Fan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
por: Nichols, Daniel, et al.
Publicado: (2025)
por: Nichols, Daniel, et al.
Publicado: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
por: Bernaschi, Massimo, et al.
Publicado: (2025)
por: Bernaschi, Massimo, et al.
Publicado: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
por: Llorente-Saguer, Isaac
Publicado: (2026)
por: Llorente-Saguer, Isaac
Publicado: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
por: Wang, Zirui, et al.
Publicado: (2026)
por: Wang, Zirui, et al.
Publicado: (2026)
MPI Implementation Profiling for Better Application Performance
por: Shipley, Riley, et al.
Publicado: (2024)
por: Shipley, Riley, et al.
Publicado: (2024)
Performance measurements of modern Fortran MPI applications with Score-P
por: Corbin, Gregor
Publicado: (2025)
por: Corbin, Gregor
Publicado: (2025)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
por: Tariq, Syed Salauddin Mohammad, et al.
Publicado: (2024)
por: Tariq, Syed Salauddin Mohammad, et al.
Publicado: (2024)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
por: Ataie, Ehsan, et al.
Publicado: (2026)
por: Ataie, Ehsan, et al.
Publicado: (2026)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
por: Ruiz, Christian Guzman, et al.
Publicado: (2024)
por: Ruiz, Christian Guzman, et al.
Publicado: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
por: Panova, Elena, et al.
Publicado: (2022)
por: Panova, Elena, et al.
Publicado: (2022)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
por: Henning, Sören, et al.
Publicado: (2025)
por: Henning, Sören, et al.
Publicado: (2025)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
por: Wen, Zhongzhen, et al.
Publicado: (2025)
por: Wen, Zhongzhen, et al.
Publicado: (2025)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
por: Heidari, Sina, et al.
Publicado: (2026)
por: Heidari, Sina, et al.
Publicado: (2026)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
por: Tuteja, Keshvi, et al.
Publicado: (2025)
por: Tuteja, Keshvi, et al.
Publicado: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
por: Li, Junjie
Publicado: (2024)
por: Li, Junjie
Publicado: (2024)
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
por: Chu, Ruifan, et al.
Publicado: (2025)
por: Chu, Ruifan, et al.
Publicado: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
por: Vogel, Adriano, et al.
Publicado: (2024)
por: Vogel, Adriano, et al.
Publicado: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
por: Bisbas, George, et al.
Publicado: (2023)
por: Bisbas, George, et al.
Publicado: (2023)
Where Should I Deploy My Contracts? A Practical Experience Report
por: Lazăr, Cătălina, et al.
Publicado: (2025)
por: Lazăr, Cătălina, et al.
Publicado: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
por: Woller, Fabian, et al.
Publicado: (2025)
por: Woller, Fabian, et al.
Publicado: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Should I Run My Cloud Benchmark on Black Friday?
por: Henning, Sören, et al.
Publicado: (2025)
por: Henning, Sören, et al.
Publicado: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
por: Tu, Jiqun, et al.
Publicado: (2026)
por: Tu, Jiqun, et al.
Publicado: (2026)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
por: Wen, Zhongzhen, et al.
Publicado: (2026)
por: Wen, Zhongzhen, et al.
Publicado: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
por: Siavashi, Mohammad, et al.
Publicado: (2026)
por: Siavashi, Mohammad, et al.
Publicado: (2026)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
por: Nichols, Daniel, et al.
Publicado: (2025)
por: Nichols, Daniel, et al.
Publicado: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
por: Lei, Yiran, et al.
Publicado: (2026)
por: Lei, Yiran, et al.
Publicado: (2026)
On the Partitioning of GPU Power among Multi-Instances
por: Vamja, Tirth, et al.
Publicado: (2025)
por: Vamja, Tirth, et al.
Publicado: (2025)
Ejemplares similares
-
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
por: Nichols, Daniel, et al.
Publicado: (2025) -
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026) -
On the energy efficiency of sparse matrix computations on multi-GPU clusters
por: Bernaschi, Massimo, et al.
Publicado: (2025) -
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
por: Llorente-Saguer, Isaac
Publicado: (2026) -
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)