FedTLU: Federated Learning with Targeted Layer Updates

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Jong-Ik, Joe-Wong, Carlee
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916583739228160
author Park, Jong-Ik
Joe-Wong, Carlee
author_facet Park, Jong-Ik
Joe-Wong, Carlee
contents Federated learning (FL) addresses privacy concerns in training language models by enabling multiple clients to contribute to the training, without sending their data to others. However, non-IID (identically and independently distributed) data across clients often limits FL's performance. This issue is especially challenging during model fine-tuning, as noise due to variations in clients' data distributions can harm model convergence near stationary points. This paper proposes a targeted layer update strategy for fine-tuning in FL. Instead of randomly updating layers of the language model, as often done in practice, we use a scoring mechanism to identify and update the most critical layers, avoiding excessively noisy or even poisoned updates by freezing the parameters in other layers. We show in extensive experiments that our method improves convergence and performance in non-IID settings, offering a more efficient approach to fine-tuning federated language models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FedTLU: Federated Learning with Targeted Layer Updates
Park, Jong-Ik
Joe-Wong, Carlee
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Federated learning (FL) addresses privacy concerns in training language models by enabling multiple clients to contribute to the training, without sending their data to others. However, non-IID (identically and independently distributed) data across clients often limits FL's performance. This issue is especially challenging during model fine-tuning, as noise due to variations in clients' data distributions can harm model convergence near stationary points. This paper proposes a targeted layer update strategy for fine-tuning in FL. Instead of randomly updating layers of the language model, as often done in practice, we use a scoring mechanism to identify and update the most critical layers, avoiding excessively noisy or even poisoned updates by freezing the parameters in other layers. We show in extensive experiments that our method improves convergence and performance in non-IID settings, offering a more efficient approach to fine-tuning federated language models.
title FedTLU: Federated Learning with Targeted Layer Updates
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2412.17692