Toward Co-adapting Machine Learning Job Shape and Cluster Topology

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Shawn Shuoshuo, Arfeen, Daiyaan, Yu, Minlan, Steenkiste, Peter, Seshan, Srinivasan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912629702787072
author Chen, Shawn Shuoshuo
Arfeen, Daiyaan
Yu, Minlan
Steenkiste, Peter
Seshan, Srinivasan
author_facet Chen, Shawn Shuoshuo
Arfeen, Daiyaan
Yu, Minlan
Steenkiste, Peter
Seshan, Srinivasan
contents Allocating resources to distributed machine learning jobs in multi-tenant torus-topology clusters must meet each job's specific placement and communication requirements, which are typically described using shapes. There is an inherent tension between minimizing network contention and maximizing cluster utilization when placing various-shaped jobs. While existing schedulers typically optimize for one objective at the expense of the other, we demonstrate that both can be achieved simultaneously. Our proposed approach, RFold, adapts both job shapes and the underlying cluster topology at runtime. This is accomplished by combining two techniques: (1) identifying homomorphic job shapes that support the jobs communication needs, and (2) reconfiguring the optical circuit switch-enabled topology to support more diverse job shapes. Preliminary evaluation performed on a 4096-node torus cluster simulator indicates that RFold can improve absolute cluster utilization by 57% and reduce job completion time by up to 11x relative to existing methods
format Preprint
id arxiv_https___arxiv_org_abs_2510_03891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Co-adapting Machine Learning Job Shape and Cluster Topology
Chen, Shawn Shuoshuo
Arfeen, Daiyaan
Yu, Minlan
Steenkiste, Peter
Seshan, Srinivasan
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Allocating resources to distributed machine learning jobs in multi-tenant torus-topology clusters must meet each job's specific placement and communication requirements, which are typically described using shapes. There is an inherent tension between minimizing network contention and maximizing cluster utilization when placing various-shaped jobs. While existing schedulers typically optimize for one objective at the expense of the other, we demonstrate that both can be achieved simultaneously. Our proposed approach, RFold, adapts both job shapes and the underlying cluster topology at runtime. This is accomplished by combining two techniques: (1) identifying homomorphic job shapes that support the jobs communication needs, and (2) reconfiguring the optical circuit switch-enabled topology to support more diverse job shapes. Preliminary evaluation performed on a 4096-node torus cluster simulator indicates that RFold can improve absolute cluster utilization by 57% and reduce job completion time by up to 11x relative to existing methods
title Toward Co-adapting Machine Learning Job Shape and Cluster Topology
topic Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2510.03891