When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Basani, Advik Raj, Vivek, Siddharth Chaitra, Krishna, Advaith, Paul, Arnab K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909528328503296
author Basani, Advik Raj
Vivek, Siddharth Chaitra
Krishna, Advaith
Paul, Arnab K.
author_facet Basani, Advik Raj
Vivek, Siddharth Chaitra
Krishna, Advaith
Paul, Arnab K.
contents Distributed Machine Learning (DML) on resource-constrained edge devices holds immense potential for real-world applications. However, achieving fast convergence in DML in these heterogeneous environments remains a significant challenge. Traditional frameworks like Bulk Synchronous Parallel and Asynchronous Stochastic Parallel rely on frequent, small updates that incur substantial communication overhead and hinder convergence speed. Furthermore, these frameworks often employ static dataset sizes, neglecting the heterogeneity of edge devices and potentially leading to straggler nodes that slow down the entire training process. The straggler nodes, i.e., edge devices that take significantly longer to process their assigned data chunk, hinder the overall training speed. To address these limitations, this paper proposes Hermes, a novel probabilistic framework for efficient DML on edge devices. This framework leverages a dynamic threshold based on recent test loss behavior to identify statistically significant improvements in the model's generalization capability, hence transmitting updates only when major improvements are detected, thereby significantly reducing communication overhead. Additionally, Hermes employs dynamic dataset allocation to optimize resource utilization and prevents performance degradation caused by straggler nodes. Our evaluations on a real-world heterogeneous resource-constrained environment demonstrate that Hermes achieves faster convergence compared to state-of-the-art methods, resulting in a remarkable $13.22$x reduction in training time and a $62.1\%$ decrease in communication overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
Basani, Advik Raj
Vivek, Siddharth Chaitra
Krishna, Advaith
Paul, Arnab K.
Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
Distributed Machine Learning (DML) on resource-constrained edge devices holds immense potential for real-world applications. However, achieving fast convergence in DML in these heterogeneous environments remains a significant challenge. Traditional frameworks like Bulk Synchronous Parallel and Asynchronous Stochastic Parallel rely on frequent, small updates that incur substantial communication overhead and hinder convergence speed. Furthermore, these frameworks often employ static dataset sizes, neglecting the heterogeneity of edge devices and potentially leading to straggler nodes that slow down the entire training process. The straggler nodes, i.e., edge devices that take significantly longer to process their assigned data chunk, hinder the overall training speed. To address these limitations, this paper proposes Hermes, a novel probabilistic framework for efficient DML on edge devices. This framework leverages a dynamic threshold based on recent test loss behavior to identify statistically significant improvements in the model's generalization capability, hence transmitting updates only when major improvements are detected, thereby significantly reducing communication overhead. Additionally, Hermes employs dynamic dataset allocation to optimize resource utilization and prevents performance degradation caused by straggler nodes. Our evaluations on a real-world heterogeneous resource-constrained environment demonstrate that Hermes achieves faster convergence compared to state-of-the-art methods, resulting in a remarkable $13.22$x reduction in training time and a $62.1\%$ decrease in communication overhead.
title When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
topic Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
url https://arxiv.org/abs/2410.20495