Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912354670739456 |
|---|---|
| author | Liu, Xinyi Wang, Yujie Zhu, Shenhan Fu, Fangcheng Liu, Qingshuo Lin, Guangming Cui, Bin |
| author_facet | Liu, Xinyi Wang, Yujie Zhu, Shenhan Fu, Fangcheng Liu, Qingshuo Lin, Guangming Cui, Bin |
| contents | Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically identifying the most efficient hybrid strategy, incorporating data, tensor, pipeline, sharded data, and sequence parallelism, along with recomputation. The system's architecture includes a profiler for hardware and model analysis, a search engine for strategy optimization using decision trees and dynamic programming, and a runtime for executing these strategies efficiently. Benchmarking on various clusters demonstrates Galvatron's superior throughput compared to existing frameworks. This open-source system offers user-friendly interfaces and comprehensive documentation, making complex distributed training accessible and efficient. The source code of Galvatron is available at https://github.com/PKU-DAIR/Hetu-Galvatron. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_21411 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Galvatron: An Automatic Distributed System for Efficient Foundation Model Training Liu, Xinyi Wang, Yujie Zhu, Shenhan Fu, Fangcheng Liu, Qingshuo Lin, Guangming Cui, Bin Distributed, Parallel, and Cluster Computing Artificial Intelligence Machine Learning Galvatron is a distributed system for efficiently training large-scale Foundation Models. It overcomes the complexities of selecting optimal parallelism strategies by automatically identifying the most efficient hybrid strategy, incorporating data, tensor, pipeline, sharded data, and sequence parallelism, along with recomputation. The system's architecture includes a profiler for hardware and model analysis, a search engine for strategy optimization using decision trees and dynamic programming, and a runtime for executing these strategies efficiently. Benchmarking on various clusters demonstrates Galvatron's superior throughput compared to existing frameworks. This open-source system offers user-friendly interfaces and comprehensive documentation, making complex distributed training accessible and efficient. The source code of Galvatron is available at https://github.com/PKU-DAIR/Hetu-Galvatron. |
| title | Galvatron: An Automatic Distributed System for Efficient Foundation Model Training |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2504.21411 |