Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Charles, Zachary, Teston, Gabriel, Dery, Lucio, Rush, Keith, Fallen, Nova, Garrett, Zachary, Szlam, Arthur, Douillard, Arthur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Federated Automatic Differentiation
von: Rush, Keith, et al.
Veröffentlicht: (2023)
von: Rush, Keith, et al.
Veröffentlicht: (2023)
DrJAX: Scalable and Differentiable MapReduce Primitives in JAX
von: Rush, Keith, et al.
Veröffentlicht: (2024)
von: Rush, Keith, et al.
Veröffentlicht: (2024)
What happens when nanochat meets DiLoCo?
von: Acker, Alexander, et al.
Veröffentlicht: (2025)
von: Acker, Alexander, et al.
Veröffentlicht: (2025)
Decoupled DiLoCo for Resilient Distributed Pre-training
von: Douillard, Arthur, et al.
Veröffentlicht: (2026)
von: Douillard, Arthur, et al.
Veröffentlicht: (2026)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
von: Douillard, Arthur, et al.
Veröffentlicht: (2025)
von: Douillard, Arthur, et al.
Veröffentlicht: (2025)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
von: Kokolis, Apostolos, et al.
Veröffentlicht: (2024)
von: Kokolis, Apostolos, et al.
Veröffentlicht: (2024)
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
von: Jaghouar, Sami, et al.
Veröffentlicht: (2024)
von: Jaghouar, Sami, et al.
Veröffentlicht: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
Optimizing Distributed Training Approaches for Scaling Neural Networks
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025)
von: Tan, Xin, et al.
Veröffentlicht: (2025)
Scalable Analysis of Urban Scaling Laws: Leveraging Cloud Computing to Analyze 21,280 Global Cities
von: Li, Zhenhui, et al.
Veröffentlicht: (2024)
von: Li, Zhenhui, et al.
Veröffentlicht: (2024)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
Automated, Reliable, and Efficient Continental-Scale Replication of 7.3 Petabytes of Climate Simulation Data: A Case Study
von: Lacinski, Lukasz, et al.
Veröffentlicht: (2024)
von: Lacinski, Lukasz, et al.
Veröffentlicht: (2024)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
von: Jamil, Hasibul, et al.
Veröffentlicht: (2025)
von: Jamil, Hasibul, et al.
Veröffentlicht: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
HPX -- An open source C++ Standard Library for Parallelism and Concurrency
von: Heller, Thomas, et al.
Veröffentlicht: (2023)
von: Heller, Thomas, et al.
Veröffentlicht: (2023)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models
von: Rajgopal, Ajay Navilarekal, et al.
Veröffentlicht: (2026)
von: Rajgopal, Ajay Navilarekal, et al.
Veröffentlicht: (2026)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
FAIR Ecosystems for Science at Scale
von: Wilkinson, Sean R., et al.
Veröffentlicht: (2025)
von: Wilkinson, Sean R., et al.
Veröffentlicht: (2025)
Scaling MPI Applications on Aurora
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Scaling Real-Time Traffic Analytics on Edge-Cloud Fabrics for City-Scale Camera Networks
von: Sharma, Akash, et al.
Veröffentlicht: (2026)
von: Sharma, Akash, et al.
Veröffentlicht: (2026)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
von: Asif, Abdullah Al, et al.
Veröffentlicht: (2026)
von: Asif, Abdullah Al, et al.
Veröffentlicht: (2026)
Scaling atomic ordering in shared memory
von: Martignetti, Lorenzo, et al.
Veröffentlicht: (2025)
von: Martignetti, Lorenzo, et al.
Veröffentlicht: (2025)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
von: Fernandez, Jared, et al.
Veröffentlicht: (2024)
von: Fernandez, Jared, et al.
Veröffentlicht: (2024)
Adaptive Self-Organization in Anonymous Dynamic Networks
von: Parzych, Garrett, et al.
Veröffentlicht: (2026)
von: Parzych, Garrett, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Federated Automatic Differentiation
von: Rush, Keith, et al.
Veröffentlicht: (2023) -
DrJAX: Scalable and Differentiable MapReduce Primitives in JAX
von: Rush, Keith, et al.
Veröffentlicht: (2024) -
What happens when nanochat meets DiLoCo?
von: Acker, Alexander, et al.
Veröffentlicht: (2025) -
Decoupled DiLoCo for Resilient Distributed Pre-training
von: Douillard, Arthur, et al.
Veröffentlicht: (2026) -
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)