Composing Distributed Computations Through Task and Kernel Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yadav, Rohan, Sundram, Shiv, Lee, Wonchan, Garland, Michael, Bauer, Michael, Aiken, Alex, Kjolstad, Fredrik
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912157480779776
author Yadav, Rohan
Sundram, Shiv
Lee, Wonchan
Garland, Michael
Bauer, Michael
Aiken, Alex
Kjolstad, Fredrik
author_facet Yadav, Rohan
Sundram, Shiv
Lee, Wonchan
Garland, Michael
Bauer, Michael
Aiken, Alex
Kjolstad, Fredrik
contents We introduce Diffuse, a system that dynamically performs task and kernel fusion in distributed, task-based runtime systems. The key component of Diffuse is an intermediate representation of distributed computation that enables the necessary analyses for the fusion of distributed tasks to be performed in a scalable manner. We pair task fusion with a JIT compiler to fuse together the kernels within fused tasks. We show empirically that Diffuse's intermediate representation is general enough to be a target for two real-world, task-based libraries (cuNumeric and Legate Sparse), letting Diffuse find optimization opportunities across function and library boundaries. Diffuse accelerates unmodified applications developed by composing task-based libraries by 1.86x on average (geo-mean), and by between 0.93x--10.7x on up to 128 GPUs. Diffuse also finds optimization opportunities missed by the original application developers, enabling high-level Python programs to match or exceed the performance of an explicitly parallel MPI library.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18109
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Composing Distributed Computations Through Task and Kernel Fusion
Yadav, Rohan
Sundram, Shiv
Lee, Wonchan
Garland, Michael
Bauer, Michael
Aiken, Alex
Kjolstad, Fredrik
Distributed, Parallel, and Cluster Computing
We introduce Diffuse, a system that dynamically performs task and kernel fusion in distributed, task-based runtime systems. The key component of Diffuse is an intermediate representation of distributed computation that enables the necessary analyses for the fusion of distributed tasks to be performed in a scalable manner. We pair task fusion with a JIT compiler to fuse together the kernels within fused tasks. We show empirically that Diffuse's intermediate representation is general enough to be a target for two real-world, task-based libraries (cuNumeric and Legate Sparse), letting Diffuse find optimization opportunities across function and library boundaries. Diffuse accelerates unmodified applications developed by composing task-based libraries by 1.86x on average (geo-mean), and by between 0.93x--10.7x on up to 128 GPUs. Diffuse also finds optimization opportunities missed by the original application developers, enabling high-level Python programs to match or exceed the performance of an explicitly parallel MPI library.
title Composing Distributed Computations Through Task and Kernel Fusion
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2406.18109