POPri: Private Federated Learning using Preference-Optimized Synthetic Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hou, Charlie, Wang, Mei-Yu, Zhu, Yige, Lazar, Daniel, Fanti, Giulia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912542897471488
author Hou, Charlie
Wang, Mei-Yu
Zhu, Yige
Lazar, Daniel
Fanti, Giulia
author_facet Hou, Charlie
Wang, Mei-Yu
Zhu, Yige
Lazar, Daniel
Fanti, Giulia
contents In practical settings, differentially private Federated learning (DP-FL) is the dominant method for training models from private, on-device client data. Recent work has suggested that DP-FL may be enhanced or outperformed by methods that use DP synthetic data (Wu et al., 2024; Hou et al., 2024). The primary algorithms for generating DP synthetic data for FL applications require careful prompt engineering based on public information and/or iterative private client feedback. Our key insight is that the private client feedback collected by prior DP synthetic data methods (Hou et al., 2024; Xie et al., 2024) can be viewed as an RL (reinforcement learning) reward. Our algorithm, Policy Optimization for Private Data (POPri) harnesses client feedback using policy optimization algorithms such as Direct Preference Optimization (DPO) to fine-tune LLMs to generate high-quality DP synthetic data. To evaluate POPri, we release LargeFedBench, a new federated text benchmark for uncontaminated LLM evaluations on federated client data. POPri substantially improves the utility of DP synthetic data relative to prior work on LargeFedBench datasets and an existing benchmark from Xie et al. (2024). POPri closes the gap between next-token prediction accuracy in the fully-private and non-private settings by up to 58%, compared to 28% for prior synthetic data methods, and 3% for state-of-the-art DP federated learning methods. The code and data are available at https://github.com/meiyuw/POPri.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16438
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POPri: Private Federated Learning using Preference-Optimized Synthetic Data
Hou, Charlie
Wang, Mei-Yu
Zhu, Yige
Lazar, Daniel
Fanti, Giulia
Machine Learning
Artificial Intelligence
Cryptography and Security
Distributed, Parallel, and Cluster Computing
In practical settings, differentially private Federated learning (DP-FL) is the dominant method for training models from private, on-device client data. Recent work has suggested that DP-FL may be enhanced or outperformed by methods that use DP synthetic data (Wu et al., 2024; Hou et al., 2024). The primary algorithms for generating DP synthetic data for FL applications require careful prompt engineering based on public information and/or iterative private client feedback. Our key insight is that the private client feedback collected by prior DP synthetic data methods (Hou et al., 2024; Xie et al., 2024) can be viewed as an RL (reinforcement learning) reward. Our algorithm, Policy Optimization for Private Data (POPri) harnesses client feedback using policy optimization algorithms such as Direct Preference Optimization (DPO) to fine-tune LLMs to generate high-quality DP synthetic data. To evaluate POPri, we release LargeFedBench, a new federated text benchmark for uncontaminated LLM evaluations on federated client data. POPri substantially improves the utility of DP synthetic data relative to prior work on LargeFedBench datasets and an existing benchmark from Xie et al. (2024). POPri closes the gap between next-token prediction accuracy in the fully-private and non-private settings by up to 58%, compared to 28% for prior synthetic data methods, and 3% for state-of-the-art DP federated learning methods. The code and data are available at https://github.com/meiyuw/POPri.
title POPri: Private Federated Learning using Preference-Optimized Synthetic Data
topic Machine Learning
Artificial Intelligence
Cryptography and Security
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.16438