CUCo: An Agentic Framework for Compute and Communication Co-design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Bodun, Varshan V, Yoga Sri, Agarwal, Saurabh, Akella, Aditya
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908861377544192
author Hu, Bodun
Varshan V, Yoga Sri
Agarwal, Saurabh
Akella, Aditya
author_facet Hu, Bodun
Varshan V, Yoga Sri
Agarwal, Saurabh
Akella, Aditya
contents Custom CUDA kernel development is essential for maximizing GPU utilization in large-scale distributed LLM training and inference, yet manually writing kernels that jointly leverage both computation and communication remains a labor-intensive and error-prone process. Prior work on kernel optimization has focused almost exclusively on computation, leaving communication kernels largely untouched even though they constitute a significant share of total execution time. We introduce CUCo, a training-free agent-driven workflow that automatically generates high-performance CUDA kernels that jointly orchestrate computation and communication. By co-optimizing these traditionally disjoint components, CUCo unlocks new optimization opportunities unavailable to existing approaches, outperforming state-of-the-art baselines and reducing end-to-end latency by up to $1.57\times$.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02376
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CUCo: An Agentic Framework for Compute and Communication Co-design
Hu, Bodun
Varshan V, Yoga Sri
Agarwal, Saurabh
Akella, Aditya
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Machine Learning
Multiagent Systems
Custom CUDA kernel development is essential for maximizing GPU utilization in large-scale distributed LLM training and inference, yet manually writing kernels that jointly leverage both computation and communication remains a labor-intensive and error-prone process. Prior work on kernel optimization has focused almost exclusively on computation, leaving communication kernels largely untouched even though they constitute a significant share of total execution time. We introduce CUCo, a training-free agent-driven workflow that automatically generates high-performance CUDA kernels that jointly orchestrate computation and communication. By co-optimizing these traditionally disjoint components, CUCo unlocks new optimization opportunities unavailable to existing approaches, outperforming state-of-the-art baselines and reducing end-to-end latency by up to $1.57\times$.
title CUCo: An Agentic Framework for Compute and Communication Co-design
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2603.02376