FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Kewen, Yi, Liping, Zhao, Zhiming, Qi, Zhuang, Yu, Han, Hu, Qinghua
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912975871279104
author Zhu, Kewen
Yi, Liping
Zhao, Zhiming
Qi, Zhuang
Yu, Han
Hu, Qinghua
author_facet Zhu, Kewen
Yi, Liping
Zhao, Zhiming
Qi, Zhuang
Yu, Han
Hu, Qinghua
contents Aligning large language models (LLMs) with human preferences in federated learning (FL) is challenging due to decentralized, privacy-sensitive, and highly non-IID preference data. Direct Preference Optimization (DPO) offers an efficient alternative to reinforcement learning with human feedback (RLHF), but its direct application in FL suffers from severe performance degradation under non-IID data and limited generalization of implicit rewards. To bridge this gap, we propose FedPDPO (Federated Personalized Direct Preference Optimization), a personalized federated framework for preference alignment of LLMs. It adopts a parameter-efficient fine-tuning architecture where each client maintains a frozen pretrained LLM backbone augmented with a Low-Rank Adaptation (LoRA) adapter, enabling communication-efficient aggregation. To address non-IID heterogeneity, we devise (1) the globally shared LoRA adapter with the personalized client-specific LLM head. Moreover, we introduce (2) a personalized DPO training strategy with a client-specific explicit reward head to complement implicit rewards and further alleviate non-IID heterogeneity, and (3) a bottleneck adapter to balance global and local features. We provide theoretical analysis establishing the probabilistic foundation and soundness. Extensive experiments on multiple preference datasets demonstrate state-of-the-art performance, achieving up to 4.80% average accuracy improvements in federated intra-domain and cross-domain settings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19741
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
Zhu, Kewen
Yi, Liping
Zhao, Zhiming
Qi, Zhuang
Yu, Han
Hu, Qinghua
Machine Learning
Computation and Language
Aligning large language models (LLMs) with human preferences in federated learning (FL) is challenging due to decentralized, privacy-sensitive, and highly non-IID preference data. Direct Preference Optimization (DPO) offers an efficient alternative to reinforcement learning with human feedback (RLHF), but its direct application in FL suffers from severe performance degradation under non-IID data and limited generalization of implicit rewards. To bridge this gap, we propose FedPDPO (Federated Personalized Direct Preference Optimization), a personalized federated framework for preference alignment of LLMs. It adopts a parameter-efficient fine-tuning architecture where each client maintains a frozen pretrained LLM backbone augmented with a Low-Rank Adaptation (LoRA) adapter, enabling communication-efficient aggregation. To address non-IID heterogeneity, we devise (1) the globally shared LoRA adapter with the personalized client-specific LLM head. Moreover, we introduce (2) a personalized DPO training strategy with a client-specific explicit reward head to complement implicit rewards and further alleviate non-IID heterogeneity, and (3) a bottleneck adapter to balance global and local features. We provide theoretical analysis establishing the probabilistic foundation and soundness. Extensive experiments on multiple preference datasets demonstrate state-of-the-art performance, achieving up to 4.80% average accuracy improvements in federated intra-domain and cross-domain settings.
title FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2603.19741