Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xin, Jihao, Canini, Marco, Richtárik, Peter, Horváth, Samuel
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915414992224256
author Xin, Jihao
Canini, Marco
Richtárik, Peter
Horváth, Samuel
author_facet Xin, Jihao
Canini, Marco
Richtárik, Peter
Horváth, Samuel
contents Distributed training enables large-scale deep learning, but suffers from high communication overhead, especially as models and datasets grow. Gradient compression, particularly quantization, is a promising approach to mitigate this bottleneck. However, existing quantization schemes are often incompatible with Allreduce, the dominant communication primitive in distributed deep learning, and many prior solutions rely on heuristics without theoretical guarantees. We introduce Global-QSGD, an Allreduce-compatible gradient quantization method that leverages global norm scaling to reduce communication overhead while preserving accuracy. Global-QSGD is backed by rigorous theoretical analysis, extending standard unbiased compressor frameworks to establish formal convergence guarantees. Additionally, we develop a performance model to evaluate its impact across different hardware configurations. Extensive experiments on NVLink, PCIe, and large-scale cloud environments show that Global-QSGD accelerates distributed training by up to 3.51% over baseline quantization methods, making it a practical and efficient solution for large-scale deep learning workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2305_18627
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
Xin, Jihao
Canini, Marco
Richtárik, Peter
Horváth, Samuel
Machine Learning
Distributed, Parallel, and Cluster Computing
I.2.11
Distributed training enables large-scale deep learning, but suffers from high communication overhead, especially as models and datasets grow. Gradient compression, particularly quantization, is a promising approach to mitigate this bottleneck. However, existing quantization schemes are often incompatible with Allreduce, the dominant communication primitive in distributed deep learning, and many prior solutions rely on heuristics without theoretical guarantees. We introduce Global-QSGD, an Allreduce-compatible gradient quantization method that leverages global norm scaling to reduce communication overhead while preserving accuracy. Global-QSGD is backed by rigorous theoretical analysis, extending standard unbiased compressor frameworks to establish formal convergence guarantees. Additionally, we develop a performance model to evaluate its impact across different hardware configurations. Extensive experiments on NVLink, PCIe, and large-scale cloud environments show that Global-QSGD accelerates distributed training by up to 3.51% over baseline quantization methods, making it a practical and efficient solution for large-scale deep learning workloads.
title Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
topic Machine Learning
Distributed, Parallel, and Cluster Computing
I.2.11
url https://arxiv.org/abs/2305.18627