NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Irene, Venkata, Vishnu Varma, Krishnamurthy, Arvind, Mahajan, Divya
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916044267847680
author Wang, Irene
Venkata, Vishnu Varma
Krishnamurthy, Arvind
Mahajan, Divya
author_facet Wang, Irene
Venkata, Vishnu Varma
Krishnamurthy, Arvind
Mahajan, Divya
contents The growing scale of deep learning demands distributed training frameworks that jointly reason about parallelism, memory, and network topology. Prior works often rely on heuristic or topology-agnostic search, handling communication and memory separately. Without per-device memory awareness, these methods typically ensure feasibility post hoc by sharding parameters and activations across many devices, increasing synchronization, inflating communication, and underutilizing compute-limiting scalability and efficiency on real datacenter networks. We present NEST, a network-, compute-, and memory-aware device placement framework that unifies model parallelism, topology modeling, and memory feasibility via structured dynamic programming. NEST's DP operates on operator graphs with tensor and expert parallel configurations, explicit allreduce latencies across hierarchical or arbitrary networks, and memory/compute profiles. By factoring parallelism across tensor, pipeline, data, and expert dimensions, NEST defines a principled search space for hybrid strategies while jointly optimizing co-location, network latency, and memory feasibility. Evaluations across diverse hardware and networks show NEST achieves up to 2.43 times higher throughput, better memory efficiency, and improved scalability over state-of-the-art baselines, providing a foundation for co-designing parallelization strategies and datacenter interconnects for next-generation AI infrastructure. The source code of NEST is available at: https://github.com/scai-tech/Nest
format Preprint
id arxiv_https___arxiv_org_abs_2603_06798
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
Wang, Irene
Venkata, Vishnu Varma
Krishnamurthy, Arvind
Mahajan, Divya
Machine Learning
Distributed, Parallel, and Cluster Computing
The growing scale of deep learning demands distributed training frameworks that jointly reason about parallelism, memory, and network topology. Prior works often rely on heuristic or topology-agnostic search, handling communication and memory separately. Without per-device memory awareness, these methods typically ensure feasibility post hoc by sharding parameters and activations across many devices, increasing synchronization, inflating communication, and underutilizing compute-limiting scalability and efficiency on real datacenter networks. We present NEST, a network-, compute-, and memory-aware device placement framework that unifies model parallelism, topology modeling, and memory feasibility via structured dynamic programming. NEST's DP operates on operator graphs with tensor and expert parallel configurations, explicit allreduce latencies across hierarchical or arbitrary networks, and memory/compute profiles. By factoring parallelism across tensor, pipeline, data, and expert dimensions, NEST defines a principled search space for hybrid strategies while jointly optimizing co-location, network latency, and memory feasibility. Evaluations across diverse hardware and networks show NEST achieves up to 2.43 times higher throughput, better memory efficiency, and improved scalability over state-of-the-art baselines, providing a foundation for co-designing parallelization strategies and datacenter interconnects for next-generation AI infrastructure. The source code of NEST is available at: https://github.com/scai-tech/Nest
title NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2603.06798