TOAST: Fast and scalable auto-partitioning based on principled static analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alabed, Sami, Grewe, Dominik, Rink, Norman Alexander, Samsikova, Masha, Sitdikov, Timur, Swietlik, Agnieszka, Vytiniotis, Dimitrios, Belov, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908499944931328
author Alabed, Sami
Grewe, Dominik
Rink, Norman Alexander
Samsikova, Masha
Sitdikov, Timur
Swietlik, Agnieszka
Vytiniotis, Dimitrios
Belov, Daniel
author_facet Alabed, Sami
Grewe, Dominik
Rink, Norman Alexander
Samsikova, Masha
Sitdikov, Timur
Swietlik, Agnieszka
Vytiniotis, Dimitrios
Belov, Daniel
contents Partitioning large machine learning models across distributed accelerator systems is a complex process, requiring a series of interdependent decisions that are further complicated by internal sharding ambiguities. Consequently, existing auto-partitioners often suffer from out-of-memory errors or are prohibitively slow when exploring the exponentially large space of possible partitionings. To mitigate this, they artificially restrict the search space, but this approach frequently yields infeasible solutions that violate device memory constraints or lead to sub-optimal performance. We propose a system that combines a novel static compiler analysis with a Monte Carlo Tree Search. Our analysis constructs an efficient decision space by identifying (i) tensor dimensions requiring identical sharding, and (ii) partitioning "conflicts" that require resolution. Our system significantly outperforms state-of-the-art industrial methods across diverse hardware platforms and model architectures, discovering previously unknown, superior solutions, and the process is fully automated even for complex and large models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TOAST: Fast and scalable auto-partitioning based on principled static analysis
Alabed, Sami
Grewe, Dominik
Rink, Norman Alexander
Samsikova, Masha
Sitdikov, Timur
Swietlik, Agnieszka
Vytiniotis, Dimitrios
Belov, Daniel
Machine Learning
Distributed, Parallel, and Cluster Computing
Partitioning large machine learning models across distributed accelerator systems is a complex process, requiring a series of interdependent decisions that are further complicated by internal sharding ambiguities. Consequently, existing auto-partitioners often suffer from out-of-memory errors or are prohibitively slow when exploring the exponentially large space of possible partitionings. To mitigate this, they artificially restrict the search space, but this approach frequently yields infeasible solutions that violate device memory constraints or lead to sub-optimal performance. We propose a system that combines a novel static compiler analysis with a Monte Carlo Tree Search. Our analysis constructs an efficient decision space by identifying (i) tensor dimensions requiring identical sharding, and (ii) partitioning "conflicts" that require resolution. Our system significantly outperforms state-of-the-art industrial methods across diverse hardware platforms and model architectures, discovering previously unknown, superior solutions, and the process is fully automated even for complex and large models.
title TOAST: Fast and scalable auto-partitioning based on principled static analysis
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.15010