Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yebo, Li, Jingguang, Tian, Chunlin, Guo, Zhijiang, Li, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911119005712384
author Wu, Yebo
Li, Jingguang
Tian, Chunlin
Guo, Zhijiang
Li, Li
author_facet Wu, Yebo
Li, Jingguang
Tian, Chunlin
Guo, Zhijiang
Li, Li
contents Federated fine-tuning enables privacy-preserving Large Language Model (LLM) adaptation, but its high memory cost limits participation from resource-constrained devices. We propose FedPruner, an innovative federated fine-tuning paradigm that tackles this via intelligent layer pruning. FedPruner flexibly prunes the global model, creating personalized submodels based on device memory constraints. It employs a macro-micro synergistic pruning framework: a macro-level functionality-driven layer orchestration mechanism groups layers, while a micro-level importance-aware layer selection strategy prunes within groups to build device-specific submodels. We further introduce a fine-grained variant that independently prunes Multi-Head Attention and Feed-Forward Network components to precisely preserve critical architectural elements. Extensive experimental results demonstrate that FedPruner significantly outperforms state-of-the-art approaches, achieving up to a 1.98\% improvement in average model accuracy while reducing peak memory usage by 75\%.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
Wu, Yebo
Li, Jingguang
Tian, Chunlin
Guo, Zhijiang
Li, Li
Distributed, Parallel, and Cluster Computing
Federated fine-tuning enables privacy-preserving Large Language Model (LLM) adaptation, but its high memory cost limits participation from resource-constrained devices. We propose FedPruner, an innovative federated fine-tuning paradigm that tackles this via intelligent layer pruning. FedPruner flexibly prunes the global model, creating personalized submodels based on device memory constraints. It employs a macro-micro synergistic pruning framework: a macro-level functionality-driven layer orchestration mechanism groups layers, while a micro-level importance-aware layer selection strategy prunes within groups to build device-specific submodels. We further introduce a fine-grained variant that independently prunes Multi-Head Attention and Feed-Forward Network components to precisely preserve critical architectural elements. Extensive experimental results demonstrate that FedPruner significantly outperforms state-of-the-art approaches, achieving up to a 1.98\% improvement in average model accuracy while reducing peak memory usage by 75\%.
title Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.17209