FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ye, Pukang, Luo, Junwei, Dong, Xiaolei, Yang, Yunbo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918194216697856
author Ye, Pukang
Luo, Junwei
Dong, Xiaolei
Yang, Yunbo
author_facet Ye, Pukang
Luo, Junwei
Dong, Xiaolei
Yang, Yunbo
contents Data duplication within large-scale corpora often impedes large language models' (LLMs) performance and privacy. In privacy-concerned federated learning scenarios, conventional deduplication methods typically rely on trusted third parties to perform uniform deletion, risking loss of informative samples while introducing privacy vulnerabilities. To address these gaps, we propose Federated ReWeighting (FedRW), the first privacy-preserving framework, to the best of our knowledge, that performs soft deduplication via sample reweighting instead of deletion in federated LLM training, without assuming a trusted third party. At its core, FedRW proposes a secure, frequency-aware reweighting protocol through secure multi-party computation, coupled with a parallel orchestration strategy to ensure efficiency and scalability. During training, FedRW utilizes an adaptive reweighting mechanism with global sample frequencies to adjust individual loss contributions, effectively improving generalization and robustness. Empirical results demonstrate that FedRW outperforms the state-of-the-art method by achieving up to 28.78x speedup in preprocessing and approximately 11.42% improvement in perplexity, while offering enhanced security guarantees. FedRW thus establishes a new paradigm for managing duplication in federated LLM training.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
Ye, Pukang
Luo, Junwei
Dong, Xiaolei
Yang, Yunbo
Cryptography and Security
Artificial Intelligence
Data duplication within large-scale corpora often impedes large language models' (LLMs) performance and privacy. In privacy-concerned federated learning scenarios, conventional deduplication methods typically rely on trusted third parties to perform uniform deletion, risking loss of informative samples while introducing privacy vulnerabilities. To address these gaps, we propose Federated ReWeighting (FedRW), the first privacy-preserving framework, to the best of our knowledge, that performs soft deduplication via sample reweighting instead of deletion in federated LLM training, without assuming a trusted third party. At its core, FedRW proposes a secure, frequency-aware reweighting protocol through secure multi-party computation, coupled with a parallel orchestration strategy to ensure efficiency and scalability. During training, FedRW utilizes an adaptive reweighting mechanism with global sample frequencies to adjust individual loss contributions, effectively improving generalization and robustness. Empirical results demonstrate that FedRW outperforms the state-of-the-art method by achieving up to 28.78x speedup in preprocessing and approximately 11.42% improvement in perplexity, while offering enhanced security guarantees. FedRW thus establishes a new paradigm for managing duplication in federated LLM training.
title FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.07505