Eager Updates For Overlapped Communication and Computation in DiLoCo
Fuente:
arXiv
Saved in:
| Main Authors: | Kale, Satyen, Douillard, Arthur, Donchev, Yanislav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025)
by: Douillard, Arthur, et al.
Published: (2025)
DiLoCo: Distributed Low-Communication Training of Language Models
by: Douillard, Arthur, et al.
Published: (2023)
by: Douillard, Arthur, et al.
Published: (2023)
Decoupled DiLoCo for Resilient Distributed Pre-training
by: Douillard, Arthur, et al.
Published: (2026)
by: Douillard, Arthur, et al.
Published: (2026)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
DiPaCo: Distributed Path Composition
by: Douillard, Arthur, et al.
Published: (2024)
by: Douillard, Arthur, et al.
Published: (2024)
Asynchronous Local-SGD Training for Language Modeling
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
What happens when nanochat meets DiLoCo?
by: Acker, Alexander, et al.
Published: (2025)
by: Acker, Alexander, et al.
Published: (2025)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
by: Defazio, Aaron, et al.
Published: (2025)
by: Defazio, Aaron, et al.
Published: (2025)
Realizability in Semantics-Guided Synthesis Done Eagerly
by: Meyer, Roland, et al.
Published: (2024)
by: Meyer, Roland, et al.
Published: (2024)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Typed Non-determinism in Concurrent Calculi: The Eager Way
by: Heuvel, Bas van den, et al.
Published: (2024)
by: Heuvel, Bas van den, et al.
Published: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
by: Hong, Ke, et al.
Published: (2025)
by: Hong, Ke, et al.
Published: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
by: Xiao, Bin, et al.
Published: (2024)
by: Xiao, Bin, et al.
Published: (2024)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025)
by: Li, Youjie, et al.
Published: (2025)
Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo
by: Shah, Vatsal, et al.
Published: (2026)
by: Shah, Vatsal, et al.
Published: (2026)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
Faster Rates For Federated Variational Inequalities
by: Wang, Guanghui, et al.
Published: (2026)
by: Wang, Guanghui, et al.
Published: (2026)
CoDi: Conversational Distillation for Grounded Question Answering
by: Huber, Patrick, et al.
Published: (2024)
by: Huber, Patrick, et al.
Published: (2024)
DiMA: An LLM-Powered Ride-Hailing Assistant at DiDi
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
JuniperLiu at CoMeDi Shared Task: Models as Annotators in Lexical Semantics Disagreements
by: Liu, Zhu, et al.
Published: (2024)
by: Liu, Zhu, et al.
Published: (2024)
CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
by: Uziel, Roy, et al.
Published: (2026)
by: Uziel, Roy, et al.
Published: (2026)
Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension
by: Kale, Sahil, et al.
Published: (2024)
by: Kale, Sahil, et al.
Published: (2024)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
by: Sarumi, Olufunke O., et al.
Published: (2025)
by: Sarumi, Olufunke O., et al.
Published: (2025)
CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation
by: Peng, Zhongyuan, et al.
Published: (2026)
by: Peng, Zhongyuan, et al.
Published: (2026)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
Bundle Adjustment in the Eager Mode
by: Zhan, Zitong, et al.
Published: (2024)
by: Zhan, Zitong, et al.
Published: (2024)
RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
KnowRL: Teaching Language Models to Know What They Know
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
Benchmarking LLMs on the Semantic Overlap Summarization Task
by: Salvador, John, et al.
Published: (2024)
by: Salvador, John, et al.
Published: (2024)
Similar Items
-
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025) -
DiLoCo: Distributed Low-Communication Training of Language Models
by: Douillard, Arthur, et al.
Published: (2023) -
Decoupled DiLoCo for Resilient Distributed Pre-training
by: Douillard, Arthur, et al.
Published: (2026) -
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025) -
DiPaCo: Distributed Path Composition
by: Douillard, Arthur, et al.
Published: (2024)