Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yuliang, Wu, Guohao, Zhang, Shenglong, Zhang, Wei, Zhu, Qianchao, Li, Zhouyang, Wang, Chenyu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2509.26246
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916980086276096
author Liu, Yuliang
Wu, Guohao
Zhang, Shenglong
Zhang, Wei
Zhu, Qianchao
Li, Zhouyang
Wang, Chenyu
author_facet Liu, Yuliang
Wu, Guohao
Zhang, Shenglong
Zhang, Wei
Zhu, Qianchao
Li, Zhouyang
Wang, Chenyu
contents The efficient distributed training of Large Language Models (LLMs) is severely hampered by the extreme variance in context lengths. This data heterogeneity, amplified by conventional packing strategies and asymmetric forward-backward costs, leads to critical inefficiencies such as cascading workload imbalances and severe hardware underutilization. Existing solutions attempt to mitigate these challenges, but often at the expense of memory or communication efficiency. To address these challenges, we introduce SlimPack, a framework that fundamentally rethinks data packing and scheduling by decomposing samples into fine-grained slices. This slice-level decomposition immediately mitigates critical memory and communication bottlenecks by transforming large, volatile workloads into a stream of smaller, manageable units. This flexibility is then harnessed for our core innovation, Asymmetric Partitioning, which assembles balanced scheduling units uniquely optimized for the different demands of the forward and backward passes. Orchestrated by a two-phase solver and a high-fidelity simulator, SlimPack holistically resolves imbalances across all parallel dimensions. Extensive experiments demonstrate that SlimPack achieves up to a $2.8\times$ training throughput improvement over baselines, breaking the conventional trade-off by delivering both superior balance and high resource efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26246
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
Liu, Yuliang
Wu, Guohao
Zhang, Shenglong
Zhang, Wei
Zhu, Qianchao
Li, Zhouyang
Wang, Chenyu
Artificial Intelligence
The efficient distributed training of Large Language Models (LLMs) is severely hampered by the extreme variance in context lengths. This data heterogeneity, amplified by conventional packing strategies and asymmetric forward-backward costs, leads to critical inefficiencies such as cascading workload imbalances and severe hardware underutilization. Existing solutions attempt to mitigate these challenges, but often at the expense of memory or communication efficiency. To address these challenges, we introduce SlimPack, a framework that fundamentally rethinks data packing and scheduling by decomposing samples into fine-grained slices. This slice-level decomposition immediately mitigates critical memory and communication bottlenecks by transforming large, volatile workloads into a stream of smaller, manageable units. This flexibility is then harnessed for our core innovation, Asymmetric Partitioning, which assembles balanced scheduling units uniquely optimized for the different demands of the forward and backward passes. Orchestrated by a two-phase solver and a high-fidelity simulator, SlimPack holistically resolves imbalances across all parallel dimensions. Extensive experiments demonstrate that SlimPack achieves up to a $2.8\times$ training throughput improvement over baselines, breaking the conventional trade-off by delivering both superior balance and high resource efficiency.
title SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
topic Artificial Intelligence
url https://arxiv.org/abs/2509.26246