Beyond A Single AI Cluster: A Survey of Decentralized LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Haotian, Jiang, Jingyan, Lu, Rongwei, Luo, Jiajun, Song, Jiajun, Li, Bowen, Shen, Ying, Wang, Zhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
von: Song, Jiajun, et al.
Veröffentlicht: (2024)
von: Song, Jiajun, et al.
Veröffentlicht: (2024)
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
von: Yang, Ween, et al.
Veröffentlicht: (2025)
von: Yang, Ween, et al.
Veröffentlicht: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
Topology-aware Federated Learning in Edge Computing: A Comprehensive Survey
von: Wu, Jiajun, et al.
Veröffentlicht: (2023)
von: Wu, Jiajun, et al.
Veröffentlicht: (2023)
The Evolution of Decentralized Systems: From Gray's Framework to Blockchain and Beyond
von: Dong, Zhongli, et al.
Veröffentlicht: (2026)
von: Dong, Zhongli, et al.
Veröffentlicht: (2026)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
PolyLink: A Blockchain Based Decentralized Edge AI Platform for LLM Inference
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
Megha: Decentralized Global Fair Scheduling for Federated Clusters
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2021)
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2021)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
A Survey on Error-Bounded Lossy Compression for Scientific Datasets
von: Di, Sheng, et al.
Veröffentlicht: (2024)
von: Di, Sheng, et al.
Veröffentlicht: (2024)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
FlowMesh: A Service Fabric for Composable LLM Workflows
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
A Survey of Synchronization Technologies for Low-power Backscatter Communication
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
Key-Embedded Privacy for Decentralized AI in Biomedical Omics
von: Zhang, Rongyu, et al.
Veröffentlicht: (2026)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2026)
Paris: A Decentralized Trained Open-Weight Diffusion Model
von: Jiang, Zhiying, et al.
Veröffentlicht: (2025)
von: Jiang, Zhiying, et al.
Veröffentlicht: (2025)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
von: Liang, Antian, et al.
Veröffentlicht: (2025)
von: Liang, Antian, et al.
Veröffentlicht: (2025)
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
von: An, Hyeonjun, et al.
Veröffentlicht: (2026)
von: An, Hyeonjun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024) -
A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
von: Song, Jiajun, et al.
Veröffentlicht: (2024) -
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
von: Yang, Ween, et al.
Veröffentlicht: (2025) -
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024) -
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)