CUCo: An Agentic Framework for Compute and Communication Co-design
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908861377544192 |
|---|---|
| author | Hu, Bodun Varshan V, Yoga Sri Agarwal, Saurabh Akella, Aditya |
| author_facet | Hu, Bodun Varshan V, Yoga Sri Agarwal, Saurabh Akella, Aditya |
| contents | Custom CUDA kernel development is essential for maximizing GPU utilization in large-scale distributed LLM training and inference, yet manually writing kernels that jointly leverage both computation and communication remains a labor-intensive and error-prone process. Prior work on kernel optimization has focused almost exclusively on computation, leaving communication kernels largely untouched even though they constitute a significant share of total execution time. We introduce CUCo, a training-free agent-driven workflow that automatically generates high-performance CUDA kernels that jointly orchestrate computation and communication. By co-optimizing these traditionally disjoint components, CUCo unlocks new optimization opportunities unavailable to existing approaches, outperforming state-of-the-art baselines and reducing end-to-end latency by up to $1.57\times$. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_02376 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CUCo: An Agentic Framework for Compute and Communication Co-design Hu, Bodun Varshan V, Yoga Sri Agarwal, Saurabh Akella, Aditya Distributed, Parallel, and Cluster Computing Hardware Architecture Machine Learning Multiagent Systems Custom CUDA kernel development is essential for maximizing GPU utilization in large-scale distributed LLM training and inference, yet manually writing kernels that jointly leverage both computation and communication remains a labor-intensive and error-prone process. Prior work on kernel optimization has focused almost exclusively on computation, leaving communication kernels largely untouched even though they constitute a significant share of total execution time. We introduce CUCo, a training-free agent-driven workflow that automatically generates high-performance CUDA kernels that jointly orchestrate computation and communication. By co-optimizing these traditionally disjoint components, CUCo unlocks new optimization opportunities unavailable to existing approaches, outperforming state-of-the-art baselines and reducing end-to-end latency by up to $1.57\times$. |
| title | CUCo: An Agentic Framework for Compute and Communication Co-design |
| topic | Distributed, Parallel, and Cluster Computing Hardware Architecture Machine Learning Multiagent Systems |
| url | https://arxiv.org/abs/2603.02376 |