StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Ziming, Wang, Shaoyu, Cheng, Shenggan, Zhao, Zhongkai, Wang, Kai, Zhao, Xuanlei, Demmel, James, You, Yang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908562316328960
author Liu, Ziming
Wang, Shaoyu
Cheng, Shenggan
Zhao, Zhongkai
Wang, Kai
Zhao, Xuanlei
Demmel, James
You, Yang
author_facet Liu, Ziming
Wang, Shaoyu
Cheng, Shenggan
Zhao, Zhongkai
Wang, Kai
Zhao, Xuanlei
Demmel, James
You, Yang
contents Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements. Specifically, StarTrail introduces an extra parallel dimension and divides the peer-to-peer communication into sub-rings to substantially reduce communication volume and avoid bandwidth bottlenecks. Through comprehensive experiments across diverse hardware environments and on both Natural Language Processing (NLP) and Computer Vision (CV) tasks, we demonstrate that our approach significantly surpasses state-of-the-art methods that support Long sequence lengths, achieving performance improvements of up to 77.12% on GPT-style models and up to 114.33% on DiT (Diffusion Transformer) models without affecting the computations results.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
Liu, Ziming
Wang, Shaoyu
Cheng, Shenggan
Zhao, Zhongkai
Wang, Kai
Zhao, Xuanlei
Demmel, James
You, Yang
Distributed, Parallel, and Cluster Computing
Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements. Specifically, StarTrail introduces an extra parallel dimension and divides the peer-to-peer communication into sub-rings to substantially reduce communication volume and avoid bandwidth bottlenecks. Through comprehensive experiments across diverse hardware environments and on both Natural Language Processing (NLP) and Computer Vision (CV) tasks, we demonstrate that our approach significantly surpasses state-of-the-art methods that support Long sequence lengths, achieving performance improvements of up to 77.12% on GPT-style models and up to 114.33% on DiT (Diffusion Transformer) models without affecting the computations results.
title StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2407.00611