Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Zhanda, Giannoula, Christina, Andoorveedu, Muralidhar, Su, Qidong, Mangalam, Karttikeya, Zheng, Bojian, Pekhimenko, Gennady
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916661859188736
author Zhu, Zhanda
Giannoula, Christina
Andoorveedu, Muralidhar
Su, Qidong
Mangalam, Karttikeya
Zheng, Bojian
Pekhimenko, Gennady
author_facet Zhu, Zhanda
Giannoula, Christina
Andoorveedu, Muralidhar
Su, Qidong
Mangalam, Karttikeya
Zheng, Bojian
Pekhimenko, Gennady
contents Various parallelism, such as data, tensor, and pipeline parallelism, along with memory optimizations like activation checkpointing, redundancy elimination, and offloading, have been proposed to accelerate distributed training for Large Language Models. To find the best combination of these techniques, automatic distributed training systems are proposed. However, existing systems only tune a subset of optimizations, due to the lack of overlap awareness, inability to navigate the vast search space, and ignoring the inter-microbatch imbalance, leading to sub-optimal performance. To address these shortcomings, we propose Mist, a memory, overlap, and imbalance-aware automatic distributed training system that comprehensively co-optimizes all memory footprint reduction techniques alongside parallelism. Mist is based on three key ideas: (1) fine-grained overlap-centric scheduling, orchestrating optimizations in an overlapped manner, (2) symbolic-based performance analysis that predicts runtime and memory usage using symbolic expressions for fast tuning, and (3) imbalance-aware hierarchical tuning, decoupling the process into an inter-stage imbalance and overlap aware Mixed Integer Linear Programming problem and an intra-stage Dual-Objective Constrained Optimization problem, and connecting them through Pareto frontier sampling. Our evaluation results show that Mist achieves an average of 1.28$\times$ (up to 1.73$\times$) and 1.27$\times$ (up to 2.04$\times$) speedup compared to state-of-the-art manual system Megatron-LM and state-of-the-art automatic system Aceso, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19050
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
Zhu, Zhanda
Giannoula, Christina
Andoorveedu, Muralidhar
Su, Qidong
Mangalam, Karttikeya
Zheng, Bojian
Pekhimenko, Gennady
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Various parallelism, such as data, tensor, and pipeline parallelism, along with memory optimizations like activation checkpointing, redundancy elimination, and offloading, have been proposed to accelerate distributed training for Large Language Models. To find the best combination of these techniques, automatic distributed training systems are proposed. However, existing systems only tune a subset of optimizations, due to the lack of overlap awareness, inability to navigate the vast search space, and ignoring the inter-microbatch imbalance, leading to sub-optimal performance. To address these shortcomings, we propose Mist, a memory, overlap, and imbalance-aware automatic distributed training system that comprehensively co-optimizes all memory footprint reduction techniques alongside parallelism. Mist is based on three key ideas: (1) fine-grained overlap-centric scheduling, orchestrating optimizations in an overlapped manner, (2) symbolic-based performance analysis that predicts runtime and memory usage using symbolic expressions for fast tuning, and (3) imbalance-aware hierarchical tuning, decoupling the process into an inter-stage imbalance and overlap aware Mixed Integer Linear Programming problem and an intra-stage Dual-Objective Constrained Optimization problem, and connecting them through Pareto frontier sampling. Our evaluation results show that Mist achieves an average of 1.28$\times$ (up to 1.73$\times$) and 1.27$\times$ (up to 2.04$\times$) speedup compared to state-of-the-art manual system Megatron-LM and state-of-the-art automatic system Aceso, respectively.
title Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2503.19050