AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kutuzov, Nikolay, Baderko, Makar, Kulibaba, Stepan, Dzhalilov, Artem, Bobrov, Daniel, Mashtaler, Maxim, Gasnikov, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916916321320960
author Kutuzov, Nikolay
Baderko, Makar
Kulibaba, Stepan
Dzhalilov, Artem
Bobrov, Daniel
Mashtaler, Maxim
Gasnikov, Alexander
author_facet Kutuzov, Nikolay
Baderko, Makar
Kulibaba, Stepan
Dzhalilov, Artem
Bobrov, Daniel
Mashtaler, Maxim
Gasnikov, Alexander
contents Scaling distributed training of Large Language Models (LLMs) requires not only algorithmic advances but also efficient utilization of heterogeneous hardware resources. While existing methods such as DiLoCo have demonstrated promising results, they often fail to fully exploit computational clusters under dynamic workloads. To address this limitation, we propose a three-stage method that combines Multi-Instance Training (MIT), Adaptive Batched DiLoCo, and switch mode mechanism. MIT allows individual nodes to run multiple lightweight training streams with different model instances in parallel and merge them to combine knowledge, increasing throughput and reducing idle time. Adaptive Batched DiLoCo dynamically adjusts local batch sizes to balance computation and communication, substantially lowering synchronization delays. Switch mode further stabilizes training by seamlessly introducing gradient accumulation once adaptive batch sizes grow beyond hardware-friendly limits. Together, these innovations improve both convergence speed and system efficiency. We also provide a theoretical estimate of the number of communications required for the full convergence of a model trained using our method.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18182
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
Kutuzov, Nikolay
Baderko, Makar
Kulibaba, Stepan
Dzhalilov, Artem
Bobrov, Daniel
Mashtaler, Maxim
Gasnikov, Alexander
Machine Learning
Artificial Intelligence
Optimization and Control
Scaling distributed training of Large Language Models (LLMs) requires not only algorithmic advances but also efficient utilization of heterogeneous hardware resources. While existing methods such as DiLoCo have demonstrated promising results, they often fail to fully exploit computational clusters under dynamic workloads. To address this limitation, we propose a three-stage method that combines Multi-Instance Training (MIT), Adaptive Batched DiLoCo, and switch mode mechanism. MIT allows individual nodes to run multiple lightweight training streams with different model instances in parallel and merge them to combine knowledge, increasing throughput and reducing idle time. Adaptive Batched DiLoCo dynamically adjusts local batch sizes to balance computation and communication, substantially lowering synchronization delays. Switch mode further stabilizes training by seamlessly introducing gradient accumulation once adaptive batch sizes grow beyond hardware-friendly limits. Together, these innovations improve both convergence speed and system efficiency. We also provide a theoretical estimate of the number of communications required for the full convergence of a model trained using our method.
title AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2508.18182