Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shyam, Vasu, Golubeva, Anna, Anthony, Quentin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026)
von: Ma, Bin, et al.
Veröffentlicht: (2026)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
von: Dutta, Ankita, et al.
Veröffentlicht: (2025)
von: Dutta, Ankita, et al.
Veröffentlicht: (2025)
Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
Pipeline Parallelism with Controllable Memory
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
TiMePReSt: Time and Memory Efficient Pipeline Parallel DNN Training with Removed Staleness
von: Dutta, Ankita, et al.
Veröffentlicht: (2024)
von: Dutta, Ankita, et al.
Veröffentlicht: (2024)
Synergistic Tensor and Pipeline Parallelism
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
von: Liu, Man, et al.
Veröffentlicht: (2026)
von: Liu, Man, et al.
Veröffentlicht: (2026)
DawnPiper: A Memory-scablable Pipeline Parallel Training Framework
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
BlockBPE: Parallel BPE Tokenization
von: You, Amos
Veröffentlicht: (2025)
von: You, Amos
Veröffentlicht: (2025)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Chameleon: Taming Dynamic Operator Sequences for Memory-Intensive LLM Training
von: Wang, Zibo, et al.
Veröffentlicht: (2025)
von: Wang, Zibo, et al.
Veröffentlicht: (2025)
A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine
von: Li, Anran, et al.
Veröffentlicht: (2026)
von: Li, Anran, et al.
Veröffentlicht: (2026)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
von: Park, Gunho, et al.
Veröffentlicht: (2022)
von: Park, Gunho, et al.
Veröffentlicht: (2022)
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
Characterizing Communication Patterns in Distributed Large Language Model Inference
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025) -
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
von: Anthony, Quentin, et al.
Veröffentlicht: (2025) -
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
von: Huang, You-Liang, et al.
Veröffentlicht: (2026) -
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024) -
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)