Asynchronous Local-SGD Training for Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Bo, Chhaparia, Rachita, Douillard, Arthur, Kale, Satyen, Rusu, Andrei A., Shen, Jiajun, Szlam, Arthur, Ranzato, Marc'Aurelio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiLoCo: Distributed Low-Communication Training of Language Models
by: Douillard, Arthur, et al.
Published: (2023)
by: Douillard, Arthur, et al.
Published: (2023)
DiPaCo: Distributed Path Composition
by: Douillard, Arthur, et al.
Published: (2024)
by: Douillard, Arthur, et al.
Published: (2024)
Context Training with Active Information Seeking
by: Huang, Zeyu, et al.
Published: (2026)
by: Huang, Zeyu, et al.
Published: (2026)
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025)
by: Douillard, Arthur, et al.
Published: (2025)
Eager Updates For Overlapped Communication and Computation in DiLoCo
by: Kale, Satyen, et al.
Published: (2025)
by: Kale, Satyen, et al.
Published: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
Decoupled DiLoCo for Resilient Distributed Pre-training
by: Douillard, Arthur, et al.
Published: (2026)
by: Douillard, Arthur, et al.
Published: (2026)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
Deliberation in Latent Space via Differentiable Cache Augmentation
by: Liu, Luyang, et al.
Published: (2024)
by: Liu, Luyang, et al.
Published: (2024)
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
Faster Rates For Federated Variational Inequalities
by: Wang, Guanghui, et al.
Published: (2026)
by: Wang, Guanghui, et al.
Published: (2026)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Latent Space Communication via K-V Cache Alignment
by: Dery, Lucio M., et al.
Published: (2026)
by: Dery, Lucio M., et al.
Published: (2026)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
by: Wuhrmann, Arthur, et al.
Published: (2025)
by: Wuhrmann, Arthur, et al.
Published: (2025)
Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD Methods
by: Tyurin, Alexander, et al.
Published: (2025)
by: Tyurin, Alexander, et al.
Published: (2025)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
by: Yakushev, George, et al.
Published: (2025)
by: Yakushev, George, et al.
Published: (2025)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Small Language Models Improve Giants by Rewriting Their Outputs
by: Vernikos, Giorgos, et al.
Published: (2023)
by: Vernikos, Giorgos, et al.
Published: (2023)
Pacing Types for Asynchronous Stream Equations
by: Kohn, Florian, et al.
Published: (2026)
by: Kohn, Florian, et al.
Published: (2026)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Ordered Momentum for Asynchronous SGD
by: Shi, Chang-Wei, et al.
Published: (2024)
by: Shi, Chang-Wei, et al.
Published: (2024)
Pacing Types: Safe Monitoring of Asynchronous Streams
by: Kohn, Florian, et al.
Published: (2025)
by: Kohn, Florian, et al.
Published: (2025)
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
Semantic Structure of Feature Space in Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2026)
by: Kozlowski, Austin C., et al.
Published: (2026)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
by: Wang, Xiaoxuan, et al.
Published: (2023)
by: Wang, Xiaoxuan, et al.
Published: (2023)
A Survey on Training-free Alignment of Large Language Models
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
On the Thinking-Language Modeling Gap in Large Language Models
by: Liu, Chenxi, et al.
Published: (2025)
by: Liu, Chenxi, et al.
Published: (2025)
Prompting Strategies for Enabling Large Language Models to Infer Causation from Correlation
by: Sgouritsa, Eleni, et al.
Published: (2024)
by: Sgouritsa, Eleni, et al.
Published: (2024)
Linear Dynamics in the RLVR Training of Large Language Models
by: Wang, Tianle, et al.
Published: (2026)
by: Wang, Tianle, et al.
Published: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
Training Language Models to Reason Efficiently
by: Arora, Daman, et al.
Published: (2025)
by: Arora, Daman, et al.
Published: (2025)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Similar Items
-
DiLoCo: Distributed Low-Communication Training of Language Models
by: Douillard, Arthur, et al.
Published: (2023) -
DiPaCo: Distributed Path Composition
by: Douillard, Arthur, et al.
Published: (2024) -
Context Training with Active Information Seeking
by: Huang, Zeyu, et al.
Published: (2026) -
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025) -
Eager Updates For Overlapped Communication and Computation in DiLoCo
by: Kale, Satyen, et al.
Published: (2025)