InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Shiju, Wang, Yujie, Sun, Ao, Fu, Fangcheng, Zhu, Zijian, Cui, Bin, Han, Xu, Ma, Kaisheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908992339443712
author Wang, Shiju
Wang, Yujie
Sun, Ao
Fu, Fangcheng
Zhu, Zijian
Cui, Bin
Han, Xu
Ma, Kaisheng
author_facet Wang, Shiju
Wang, Yujie
Sun, Ao
Fu, Fangcheng
Zhu, Zijian
Cui, Bin
Han, Xu
Ma, Kaisheng
contents Long context training is crucial for LLM's context extension. Existing schemes, such as sequence parallelism, incur substantial communication overhead. Pipeline parallelism (PP) reduces this cost, but its effectiveness hinges on partitioning granularity. Batch-level PP employing sequence packing exhibits high memory consumption in long-context scenarios, whereas token-level PP splitting sequences into slices alleviates memory overhead but may incur hardware under-utilization. Moreover, the skewed distribution of sequence length in real-world datasets renders monolithic and static granularity PP's sub-optimal performance. In this paper, we propose 1) \textit{Elastic Pipeline Parallelism} (EPP) that orchestrates token-level PP and batch-level PP to adapt to resource and workload heterogeneity, and 2) \textit{Stage-Aware Chunk-Level Adaptive Checkpointing} that efficiently integrates gradient checkpointing with EPP. Comprehensive experiments demonstrate that InfiniPipe achieves a 1.69x speedup over state-of-the-art systems. Our code is open-sourced at https://github.com/wsjdsg/InfiniPipe-code.git.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21275
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
Wang, Shiju
Wang, Yujie
Sun, Ao
Fu, Fangcheng
Zhu, Zijian
Cui, Bin
Han, Xu
Ma, Kaisheng
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Long context training is crucial for LLM's context extension. Existing schemes, such as sequence parallelism, incur substantial communication overhead. Pipeline parallelism (PP) reduces this cost, but its effectiveness hinges on partitioning granularity. Batch-level PP employing sequence packing exhibits high memory consumption in long-context scenarios, whereas token-level PP splitting sequences into slices alleviates memory overhead but may incur hardware under-utilization. Moreover, the skewed distribution of sequence length in real-world datasets renders monolithic and static granularity PP's sub-optimal performance. In this paper, we propose 1) \textit{Elastic Pipeline Parallelism} (EPP) that orchestrates token-level PP and batch-level PP to adapt to resource and workload heterogeneity, and 2) \textit{Stage-Aware Chunk-Level Adaptive Checkpointing} that efficiently integrates gradient checkpointing with EPP. Comprehensive experiments demonstrate that InfiniPipe achieves a 1.69x speedup over state-of-the-art systems. Our code is open-sourced at https://github.com/wsjdsg/InfiniPipe-code.git.
title InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2509.21275