Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Jiangfei, Zhang, Shuo, Wang, Zerui, Jiang, Lijuan, Qu, Wenwen, Hu, Qinghao, Wang, Guoteng, Weng, Qizhen, Yan, Hang, Zhang, Xingcheng, Qiu, Xipeng, Lin, Dahua, Wen, Yonggang, Jin, Xin, Zhang, Tianwei, Sun, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
von: Wang, Zerui, et al.
Veröffentlicht: (2025)
von: Wang, Zerui, et al.
Veröffentlicht: (2025)
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
von: Zhang, Meng, et al.
Veröffentlicht: (2024)
von: Zhang, Meng, et al.
Veröffentlicht: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
von: Li, Senyao, et al.
Veröffentlicht: (2025)
von: Li, Senyao, et al.
Veröffentlicht: (2025)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
HopGNN: Boosting Distributed GNN Training Efficiency via Feature-Centric Model Migration
von: Chen, Weijian, et al.
Veröffentlicht: (2024)
von: Chen, Weijian, et al.
Veröffentlicht: (2024)
Robust LLM Training Infrastructure at ByteDance
von: Wan, Borui, et al.
Veröffentlicht: (2025)
von: Wan, Borui, et al.
Veröffentlicht: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
A Survey of Synchronization Technologies for Low-power Backscatter Communication
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
von: Wu, Du, et al.
Veröffentlicht: (2025)
von: Wu, Du, et al.
Veröffentlicht: (2025)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
von: Duan, Tao, et al.
Veröffentlicht: (2025)
von: Duan, Tao, et al.
Veröffentlicht: (2025)
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
von: Chen, Xinquan, et al.
Veröffentlicht: (2026)
von: Chen, Xinquan, et al.
Veröffentlicht: (2026)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
Incremental GNN Embedding Computation on Streaming Graphs
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
TrainMover: An Interruption-Resilient Runtime for ML Training
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
von: Wang, Zerui, et al.
Veröffentlicht: (2025) -
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024) -
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023) -
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)