AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Qiaoling, Hu, Qinghao, Wang, Guoteng, Xiong, Yingtong, Huang, Ting, Chen, Xun, Gao, Yang, Yan, Hang, Wen, Yonggang, Zhang, Tianwei, Sun, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024)
by: Gu, Diandian, et al.
Published: (2024)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
by: Bai, Huawei, et al.
Published: (2025)
by: Bai, Huawei, et al.
Published: (2025)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
by: Chen, Qiaoling, et al.
Published: (2024)
by: Chen, Qiaoling, et al.
Published: (2024)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
Characterization of Large Language Model Development in the Datacenter
by: Hu, Qinghao, et al.
Published: (2024)
by: Hu, Qinghao, et al.
Published: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
by: Gong, Yidong, et al.
Published: (2026)
by: Gong, Yidong, et al.
Published: (2026)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
by: Zhang, Meng, et al.
Published: (2024)
by: Zhang, Meng, et al.
Published: (2024)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2025)
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2025)
FeedSign: Robust Full-parameter Federated Fine-tuning of Large Models with Extremely Low Communication Overhead of One Bit
by: Cai, Zhijie, et al.
Published: (2025)
by: Cai, Zhijie, et al.
Published: (2025)
Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
Ocior: Ultra-Fast Asynchronous Leaderless Consensus with Two-Round Finality, Linear Overhead, and Adaptive Security
by: Chen, Jinyuan
Published: (2025)
by: Chen, Jinyuan
Published: (2025)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
by: Zhang, Songge, et al.
Published: (2026)
by: Zhang, Songge, et al.
Published: (2026)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
by: Chen, Fahao, et al.
Published: (2024)
by: Chen, Fahao, et al.
Published: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
by: Chen, Xiaopei, et al.
Published: (2025)
by: Chen, Xiaopei, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
Communication-Efficient Model Aggregation with Layer Divergence Feedback in Federated Learning
by: Wang, Liwei, et al.
Published: (2024)
by: Wang, Liwei, et al.
Published: (2024)
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024)
by: Xiong, Dian, et al.
Published: (2024)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
by: Strati, Foteini, et al.
Published: (2025)
by: Strati, Foteini, et al.
Published: (2025)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
by: Xu, Guanbin, et al.
Published: (2026)
by: Xu, Guanbin, et al.
Published: (2026)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Efficient Distributed MLLM Training with Cornstarch
by: Jang, Insu, et al.
Published: (2025)
by: Jang, Insu, et al.
Published: (2025)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
by: Ishikawa, Satoki, et al.
Published: (2024)
by: Ishikawa, Satoki, et al.
Published: (2024)
Modular Architecture for High-Performance and Low Overhead Data Transfers
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
by: Gan, Zhenghao, et al.
Published: (2026)
by: Gan, Zhenghao, et al.
Published: (2026)
Federated Inference for Heterogeneous LLM Communication and Collaboration
by: Chen, Zihan, et al.
Published: (2026)
by: Chen, Zihan, et al.
Published: (2026)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
by: Uhl, Tim Niklas, et al.
Published: (2024)
by: Uhl, Tim Niklas, et al.
Published: (2024)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
by: Licker, Nandor, et al.
Published: (2025)
by: Licker, Nandor, et al.
Published: (2025)
Similar Items
-
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024) -
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
by: Bai, Huawei, et al.
Published: (2025) -
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
by: Chen, Qiaoling, et al.
Published: (2024) -
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
by: Chen, Qiaoling, et al.
Published: (2025) -
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)