Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Qianren, Zhang, Qili, Hao, Hanwen, Han, Zhentao, Xu, Runhua, Jiang, Weifeng, Hu, Qi, Chen, Zhijun, Zhou, Tyler, Li, Bo, Song, Yangqiu, Dong, Jin, Li, Jianxin, Yu, Philip S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915261996597248
author Mao, Qianren
Zhang, Qili
Hao, Hanwen
Han, Zhentao
Xu, Runhua
Jiang, Weifeng
Hu, Qi
Chen, Zhijun
Zhou, Tyler
Li, Bo
Song, Yangqiu
Dong, Jin
Li, Jianxin
Yu, Philip S.
author_facet Mao, Qianren
Zhang, Qili
Hao, Hanwen
Han, Zhentao
Xu, Runhua
Jiang, Weifeng
Hu, Qi
Chen, Zhijun
Zhou, Tyler
Li, Bo
Song, Yangqiu
Dong, Jin
Li, Jianxin
Yu, Philip S.
contents Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution for enhancing the accuracy and credibility of Large Language Models (LLMs), particularly in Question & Answer tasks. This is achieved by incorporating proprietary and private data from integrated databases. However, private RAG systems face significant challenges due to the scarcity of private domain data and critical data privacy issues. These obstacles impede the deployment of private RAG systems, as developing privacy-preserving RAG systems requires a delicate balance between data security and data availability. To address these challenges, we regard federated learning (FL) as a highly promising technology for privacy-preserving RAG services. We propose a novel framework called Federated Retrieval-Augmented Generation (FedE4RAG). This framework facilitates collaborative training of client-side RAG retrieval models. The parameters of these models are aggregated and distributed on a central-server, ensuring data privacy without direct sharing of raw data. In FedE4RAG, knowledge distillation is employed for communication between the server and client models. This technique improves the generalization of local RAG retrievers during the federated learning process. Additionally, we apply homomorphic encryption within federated learning to safeguard model parameters and mitigate concerns related to data leakage. Extensive experiments conducted on the real-world dataset have validated the effectiveness of FedE4RAG. The results demonstrate that our proposed framework can markedly enhance the performance of private RAG systems while maintaining robust data privacy protection.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19101
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation
Mao, Qianren
Zhang, Qili
Hao, Hanwen
Han, Zhentao
Xu, Runhua
Jiang, Weifeng
Hu, Qi
Chen, Zhijun
Zhou, Tyler
Li, Bo
Song, Yangqiu
Dong, Jin
Li, Jianxin
Yu, Philip S.
Computation and Language
Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution for enhancing the accuracy and credibility of Large Language Models (LLMs), particularly in Question & Answer tasks. This is achieved by incorporating proprietary and private data from integrated databases. However, private RAG systems face significant challenges due to the scarcity of private domain data and critical data privacy issues. These obstacles impede the deployment of private RAG systems, as developing privacy-preserving RAG systems requires a delicate balance between data security and data availability. To address these challenges, we regard federated learning (FL) as a highly promising technology for privacy-preserving RAG services. We propose a novel framework called Federated Retrieval-Augmented Generation (FedE4RAG). This framework facilitates collaborative training of client-side RAG retrieval models. The parameters of these models are aggregated and distributed on a central-server, ensuring data privacy without direct sharing of raw data. In FedE4RAG, knowledge distillation is employed for communication between the server and client models. This technique improves the generalization of local RAG retrievers during the federated learning process. Additionally, we apply homomorphic encryption within federated learning to safeguard model parameters and mitigate concerns related to data leakage. Extensive experiments conducted on the real-world dataset have validated the effectiveness of FedE4RAG. The results demonstrate that our proposed framework can markedly enhance the performance of private RAG systems while maintaining robust data privacy protection.
title Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2504.19101