Heterogeneous Federated Learning with Splited Language Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shi, Yifan, Zhang, Yuhui, Huang, Ziyue, Yang, Xiaofeng, Shen, Li, Chen, Wei, Wang, Xueqian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913321959030784
author Shi, Yifan
Zhang, Yuhui
Huang, Ziyue
Yang, Xiaofeng
Shen, Li
Chen, Wei
Wang, Xueqian
author_facet Shi, Yifan
Zhang, Yuhui
Huang, Ziyue
Yang, Xiaofeng
Shen, Li
Chen, Wei
Wang, Xueqian
contents Federated Split Learning (FSL) is a promising distributed learning paradigm in practice, which gathers the strengths of both Federated Learning (FL) and Split Learning (SL) paradigms, to ensure model privacy while diminishing the resource overhead of each client, especially on large transformer models in a resource-constrained environment, e.g., Internet of Things (IoT). However, almost all works merely investigate the performance with simple neural network models in FSL. Despite the minor efforts focusing on incorporating Vision Transformers (ViT) as model architectures, they train ViT from scratch, thereby leading to enormous training overhead in each device with limited resources. Therefore, in this paper, we harness Pre-trained Image Transformers (PITs) as the initial model, coined FedV, to accelerate the training process and improve model robustness. Furthermore, we propose FedVZ to hinder the gradient inversion attack, especially having the capability compatible with black-box scenarios, where the gradient information is unavailable. Concretely, FedVZ approximates the server gradient by utilizing a zeroth-order (ZO) optimization, which replaces the backward propagation with just one forward process. Empirically, we are the first to provide a systematic evaluation of FSL methods with PITs in real-world datasets, different partial device participations, and heterogeneous data splits. Our experiments verify the effectiveness of our algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16050
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Heterogeneous Federated Learning with Splited Language Model
Shi, Yifan
Zhang, Yuhui
Huang, Ziyue
Yang, Xiaofeng
Shen, Li
Chen, Wei
Wang, Xueqian
Computer Vision and Pattern Recognition
Federated Split Learning (FSL) is a promising distributed learning paradigm in practice, which gathers the strengths of both Federated Learning (FL) and Split Learning (SL) paradigms, to ensure model privacy while diminishing the resource overhead of each client, especially on large transformer models in a resource-constrained environment, e.g., Internet of Things (IoT). However, almost all works merely investigate the performance with simple neural network models in FSL. Despite the minor efforts focusing on incorporating Vision Transformers (ViT) as model architectures, they train ViT from scratch, thereby leading to enormous training overhead in each device with limited resources. Therefore, in this paper, we harness Pre-trained Image Transformers (PITs) as the initial model, coined FedV, to accelerate the training process and improve model robustness. Furthermore, we propose FedVZ to hinder the gradient inversion attack, especially having the capability compatible with black-box scenarios, where the gradient information is unavailable. Concretely, FedVZ approximates the server gradient by utilizing a zeroth-order (ZO) optimization, which replaces the backward propagation with just one forward process. Empirically, we are the first to provide a systematic evaluation of FSL methods with PITs in real-world datasets, different partial device participations, and heterogeneous data splits. Our experiments verify the effectiveness of our algorithms.
title Heterogeneous Federated Learning with Splited Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16050