Astra: A Multi-Agent System for GPU Kernel Performance Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Anjiang, Sun, Tianran, Seenichamy, Yogesh, Song, Hang, Ouyang, Anne, Mirhoseini, Azalia, Wang, Ke, Aiken, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Task-Based Programming for Adaptive Mesh Refinement in Compressible Flow Simulations
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
by: Saba, Tara, et al.
Published: (2026)
by: Saba, Tara, et al.
Published: (2026)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
by: Wang, Hansheng, et al.
Published: (2025)
by: Wang, Hansheng, et al.
Published: (2025)
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
by: Wei, Anjiang, et al.
Published: (2024)
by: Wei, Anjiang, et al.
Published: (2024)
Do Large Language Models Understand Performance Optimization?
by: Cui, Bowen, et al.
Published: (2025)
by: Cui, Bowen, et al.
Published: (2025)
Multi-Grained Specifications for Distributed System Model Checking and Verification
by: Ouyang, Lingzhi, et al.
Published: (2024)
by: Ouyang, Lingzhi, et al.
Published: (2024)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
by: Singh, Abhinav, et al.
Published: (2023)
by: Singh, Abhinav, et al.
Published: (2023)
SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions
by: Lee, Cheryl, et al.
Published: (2024)
by: Lee, Cheryl, et al.
Published: (2024)
Comprehensive Review of Performance Optimization Strategies for Serverless Applications on AWS Lambda
by: Bechir, Mohamed Lemine El, et al.
Published: (2024)
by: Bechir, Mohamed Lemine El, et al.
Published: (2024)
Investigating Matrix Repartitioning to Address the Over- and Undersubscription Challenge for a GPU-based CFD Solver
by: Olenik, Gregor, et al.
Published: (2025)
by: Olenik, Gregor, et al.
Published: (2025)
MARCO: Multi-Agent Code Optimization with Real-Time Knowledge Integration for High-Performance Computing
by: Rahman, Asif, et al.
Published: (2025)
by: Rahman, Asif, et al.
Published: (2025)
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
by: Zhu, Xinguo, et al.
Published: (2025)
by: Zhu, Xinguo, et al.
Published: (2025)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
by: Matsumura, Kazuaki, et al.
Published: (2023)
by: Matsumura, Kazuaki, et al.
Published: (2023)
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
by: Zhao, Xingzhong, et al.
Published: (2026)
by: Zhao, Xingzhong, et al.
Published: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
by: Wen, Zhongzhen, et al.
Published: (2025)
by: Wen, Zhongzhen, et al.
Published: (2025)
High-Performance Star-M SVD for Big Data Compression
by: Hussain, Md Taufique, et al.
Published: (2026)
by: Hussain, Md Taufique, et al.
Published: (2026)
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
by: Timalsina, Madan, et al.
Published: (2024)
by: Timalsina, Madan, et al.
Published: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
by: Domke, Jens, et al.
Published: (2025)
by: Domke, Jens, et al.
Published: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
by: Tschand, Arya, et al.
Published: (2025)
by: Tschand, Arya, et al.
Published: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
Multi-Objective Load Balancing for Heterogeneous Edge-Based Object Detection Systems
by: Alqahtani, Daghash K., et al.
Published: (2026)
by: Alqahtani, Daghash K., et al.
Published: (2026)
Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure
by: Pagidoju, Ravi Teja
Published: (2026)
by: Pagidoju, Ravi Teja
Published: (2026)
GPU Implementations for Midsize Integer Addition and Multiplication
by: Oancea, Cosmin E., et al.
Published: (2024)
by: Oancea, Cosmin E., et al.
Published: (2024)
Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach
by: Picatto, Hernan, et al.
Published: (2024)
by: Picatto, Hernan, et al.
Published: (2024)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
by: Jangda, Abhinav, et al.
Published: (2023)
by: Jangda, Abhinav, et al.
Published: (2023)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
SPUMA: a minimally invasive approach to the GPU porting of OPENFOAM
by: Bnà, Simone, et al.
Published: (2025)
by: Bnà, Simone, et al.
Published: (2025)
Building AI Agents for Autonomous Clouds: Challenges and Design Principles
by: Shetty, Manish, et al.
Published: (2024)
by: Shetty, Manish, et al.
Published: (2024)
Similar Items
-
Task-Based Programming for Adaptive Mesh Refinement in Compressible Flow Simulations
by: Wei, Anjiang, et al.
Published: (2025) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025) -
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025) -
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)