Galvatron: An Automatic Distributed System for Efficient Foundation Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xinyi, Wang, Yujie, Zhu, Shenhan, Fu, Fangcheng, Liu, Qingshuo, Lin, Guangming, Cui, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912354670739456
author Liu, Xinyi
Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Liu, Qingshuo
Lin, Guangming
Cui, Bin
author_facet Liu, Xinyi
Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Liu, Qingshuo
Lin, Guangming
Cui, Bin
contents Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically identifying the most efficient hybrid strategy, incorporating data, tensor, pipeline, sharded data, and sequence parallelism, along with recomputation. The system's architecture includes a profiler for hardware and model analysis, a search engine for strategy optimization using decision trees and dynamic programming, and a runtime for executing these strategies efficiently. Benchmarking on various clusters demonstrates Galvatron's superior throughput compared to existing frameworks. This open-source system offers user-friendly interfaces and comprehensive documentation, making complex distributed training accessible and efficient. The source code of Galvatron is available at https://github.com/PKU-DAIR/Hetu-Galvatron.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
Liu, Xinyi
Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Liu, Qingshuo
Lin, Guangming
Cui, Bin
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically identifying the most efficient hybrid strategy, incorporating data, tensor, pipeline, sharded data, and sequence parallelism, along with recomputation. The system's architecture includes a profiler for hardware and model analysis, a search engine for strategy optimization using decision trees and dynamic programming, and a runtime for executing these strategies efficiently. Benchmarking on various clusters demonstrates Galvatron's superior throughput compared to existing frameworks. This open-source system offers user-friendly interfaces and comprehensive documentation, making complex distributed training accessible and efficient. The source code of Galvatron is available at https://github.com/PKU-DAIR/Hetu-Galvatron.
title Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.21411