Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911119005712384 |
|---|---|
| author | Wu, Yebo Li, Jingguang Tian, Chunlin Guo, Zhijiang Li, Li |
| author_facet | Wu, Yebo Li, Jingguang Tian, Chunlin Guo, Zhijiang Li, Li |
| contents | Federated fine-tuning enables privacy-preserving Large Language Model (LLM) adaptation, but its high memory cost limits participation from resource-constrained devices. We propose FedPruner, an innovative federated fine-tuning paradigm that tackles this via intelligent layer pruning. FedPruner flexibly prunes the global model, creating personalized submodels based on device memory constraints. It employs a macro-micro synergistic pruning framework: a macro-level functionality-driven layer orchestration mechanism groups layers, while a micro-level importance-aware layer selection strategy prunes within groups to build device-specific submodels. We further introduce a fine-grained variant that independently prunes Multi-Head Attention and Feed-Forward Network components to precisely preserve critical architectural elements. Extensive experimental results demonstrate that FedPruner significantly outperforms state-of-the-art approaches, achieving up to a 1.98\% improvement in average model accuracy while reducing peak memory usage by 75\%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_17209 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning Wu, Yebo Li, Jingguang Tian, Chunlin Guo, Zhijiang Li, Li Distributed, Parallel, and Cluster Computing Federated fine-tuning enables privacy-preserving Large Language Model (LLM) adaptation, but its high memory cost limits participation from resource-constrained devices. We propose FedPruner, an innovative federated fine-tuning paradigm that tackles this via intelligent layer pruning. FedPruner flexibly prunes the global model, creating personalized submodels based on device memory constraints. It employs a macro-micro synergistic pruning framework: a macro-level functionality-driven layer orchestration mechanism groups layers, while a micro-level importance-aware layer selection strategy prunes within groups to build device-specific submodels. We further introduce a fine-grained variant that independently prunes Multi-Head Attention and Feed-Forward Network components to precisely preserve critical architectural elements. Extensive experimental results demonstrate that FedPruner significantly outperforms state-of-the-art approaches, achieving up to a 1.98\% improvement in average model accuracy while reducing peak memory usage by 75\%. |
| title | Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2508.17209 |