Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models with Structured Pruning in Resource-Limited Clients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yan, Zhang, Xiao, Li, Mingyi, Xu, Guangwei, Chen, Feng, Yuan, Yuan, Zou, Yifei, Zhao, Mengying, Lu, Jianbo, Yu, Dongxiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915570306252800
author Li, Yan
Zhang, Xiao
Li, Mingyi
Xu, Guangwei
Chen, Feng
Yuan, Yuan
Zou, Yifei
Zhao, Mengying
Lu, Jianbo
Yu, Dongxiao
author_facet Li, Yan
Zhang, Xiao
Li, Mingyi
Xu, Guangwei
Chen, Feng
Yuan, Yuan
Zou, Yifei
Zhao, Mengying
Lu, Jianbo
Yu, Dongxiao
contents In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the \textit{unstructured pruning}, \textit{varying submodel architectures}, \textit{knowledge loss}, and \textit{straggler} challenges simultaneously. We propose a novel semi-asynchronous collaborative training framework, namely ${Co\text{-}S}^2{P}$, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that ${Co\text{-}S}^2{P}$ can achieve asymptotic optimal convergence rate of $O(1/\sqrt{N^*EQ})$. Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices.The experimental results demonstrate that $Co\text{-}S^2P$ improves accuracy by up to 8.8\% and resource utilization by up to 1.2$\times$ compared to state-of-the-art methods, while reducing memory consumption by approximately 22\% and training time by about 24\% on all resource-limited devices.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08457
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models with Structured Pruning in Resource-Limited Clients
Li, Yan
Zhang, Xiao
Li, Mingyi
Xu, Guangwei
Chen, Feng
Yuan, Yuan
Zou, Yifei
Zhao, Mengying
Lu, Jianbo
Yu, Dongxiao
Distributed, Parallel, and Cluster Computing
Machine Learning
In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the \textit{unstructured pruning}, \textit{varying submodel architectures}, \textit{knowledge loss}, and \textit{straggler} challenges simultaneously. We propose a novel semi-asynchronous collaborative training framework, namely ${Co\text{-}S}^2{P}$, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that ${Co\text{-}S}^2{P}$ can achieve asymptotic optimal convergence rate of $O(1/\sqrt{N^*EQ})$. Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices.The experimental results demonstrate that $Co\text{-}S^2P$ improves accuracy by up to 8.8\% and resource utilization by up to 1.2$\times$ compared to state-of-the-art methods, while reducing memory consumption by approximately 22\% and training time by about 24\% on all resource-limited devices.
title Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models with Structured Pruning in Resource-Limited Clients
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2410.08457