Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Ji, Ren, Jiaxiang, Jin, Ruoming, Zhang, Zijie, Zhou, Yang, Valduriez, Patrick, Dou, Dejing
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916444160131072
author Liu, Ji
Ren, Jiaxiang
Jin, Ruoming
Zhang, Zijie
Zhou, Yang
Valduriez, Patrick
Dou, Dejing
author_facet Liu, Ji
Ren, Jiaxiang
Jin, Ruoming
Zhang, Zijie
Zhou, Yang
Valduriez, Patrick
Dou, Dejing
contents As a promising paradigm to collaboratively train models with decentralized data, Federated Learning (FL) can be exploited to fine-tune Large Language Models (LLMs). While LLMs correspond to huge size, the scale of the training data significantly increases, which leads to tremendous amounts of computation and communication costs. The training data is generally non-Independent and Identically Distributed (non-IID), which requires adaptive data processing within each device. Although Low Rank Adaptation (LoRA) can significantly reduce the scale of parameters to update in the fine-tuning process, it still takes unaffordable time to transfer the low-rank parameters of all the layers in LLMs. In this paper, we propose a Fisher Information-based Efficient Curriculum Federated Learning framework (FibecFed) with two novel methods, i.e., adaptive federated curriculum learning and efficient sparse parameter update. First, we propose a fisher information-based method to adaptively sample data within each device to improve the effectiveness of the FL fine-tuning process. Second, we dynamically select the proper layers for global aggregation and sparse parameters for local update with LoRA so as to improve the efficiency of the FL fine-tuning process. Extensive experimental results based on 10 datasets demonstrate that FibecFed yields excellent performance (up to 45.35% in terms of accuracy) and superb fine-tuning speed (up to 98.61% faster) compared with 17 baseline approaches).
format Preprint
id arxiv_https___arxiv_org_abs_2410_00131
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models
Liu, Ji
Ren, Jiaxiang
Jin, Ruoming
Zhang, Zijie
Zhou, Yang
Valduriez, Patrick
Dou, Dejing
Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
As a promising paradigm to collaboratively train models with decentralized data, Federated Learning (FL) can be exploited to fine-tune Large Language Models (LLMs). While LLMs correspond to huge size, the scale of the training data significantly increases, which leads to tremendous amounts of computation and communication costs. The training data is generally non-Independent and Identically Distributed (non-IID), which requires adaptive data processing within each device. Although Low Rank Adaptation (LoRA) can significantly reduce the scale of parameters to update in the fine-tuning process, it still takes unaffordable time to transfer the low-rank parameters of all the layers in LLMs. In this paper, we propose a Fisher Information-based Efficient Curriculum Federated Learning framework (FibecFed) with two novel methods, i.e., adaptive federated curriculum learning and efficient sparse parameter update. First, we propose a fisher information-based method to adaptively sample data within each device to improve the effectiveness of the FL fine-tuning process. Second, we dynamically select the proper layers for global aggregation and sparse parameters for local update with LoRA so as to improve the efficiency of the FL fine-tuning process. Extensive experimental results based on 10 datasets demonstrate that FibecFed yields excellent performance (up to 45.35% in terms of accuracy) and superb fine-tuning speed (up to 98.61% faster) compared with 17 baseline approaches).
title Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2410.00131