Get More for Less in Decentralized Learning Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhasade, Akash, Kermarrec, Anne-Marie, Pires, Rafael, Sharma, Rishi, Vujasinovic, Milos, Wigger, Jeffrey
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909201315397632
author Dhasade, Akash
Kermarrec, Anne-Marie
Pires, Rafael
Sharma, Rishi
Vujasinovic, Milos
Wigger, Jeffrey
author_facet Dhasade, Akash
Kermarrec, Anne-Marie
Pires, Rafael
Sharma, Rishi
Vujasinovic, Milos
Wigger, Jeffrey
contents Decentralized learning (DL) systems have been gaining popularity because they avoid raw data sharing by communicating only model parameters, hence preserving data confidentiality. However, the large size of deep neural networks poses a significant challenge for decentralized training, since each node needs to exchange gigabytes of data, overloading the network. In this paper, we address this challenge with JWINS, a communication-efficient and fully decentralized learning system that shares only a subset of parameters through sparsification. JWINS uses wavelet transform to limit the information loss due to sparsification and a randomized communication cut-off that reduces communication usage without damaging the performance of trained models. We demonstrate empirically with 96 DL nodes on non-IID datasets that JWINS can achieve similar accuracies to full-sharing DL while sending up to 64% fewer bytes. Additionally, on low communication budgets, JWINS outperforms the state-of-the-art communication-efficient DL algorithm CHOCO-SGD by up to 4x in terms of network savings and time.
format Preprint
id arxiv_https___arxiv_org_abs_2306_04377
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Get More for Less in Decentralized Learning Systems
Dhasade, Akash
Kermarrec, Anne-Marie
Pires, Rafael
Sharma, Rishi
Vujasinovic, Milos
Wigger, Jeffrey
Distributed, Parallel, and Cluster Computing
Machine Learning
Decentralized learning (DL) systems have been gaining popularity because they avoid raw data sharing by communicating only model parameters, hence preserving data confidentiality. However, the large size of deep neural networks poses a significant challenge for decentralized training, since each node needs to exchange gigabytes of data, overloading the network. In this paper, we address this challenge with JWINS, a communication-efficient and fully decentralized learning system that shares only a subset of parameters through sparsification. JWINS uses wavelet transform to limit the information loss due to sparsification and a randomized communication cut-off that reduces communication usage without damaging the performance of trained models. We demonstrate empirically with 96 DL nodes on non-IID datasets that JWINS can achieve similar accuracies to full-sharing DL while sending up to 64% fewer bytes. Additionally, on low communication budgets, JWINS outperforms the state-of-the-art communication-efficient DL algorithm CHOCO-SGD by up to 4x in terms of network savings and time.
title Get More for Less in Decentralized Learning Systems
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2306.04377