PluralLLM: Pluralistic Alignment in LLMs via Federated Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Srewa, Mahmoud, Zhao, Tianyu, Elmalaki, Salma
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910872684724224
author Srewa, Mahmoud
Zhao, Tianyu
Elmalaki, Salma
author_facet Srewa, Mahmoud
Zhao, Tianyu
Elmalaki, Salma
contents Ensuring Large Language Models (LLMs) align with diverse human preferences while preserving privacy and fairness remains a challenge. Existing methods, such as Reinforcement Learning from Human Feedback (RLHF), rely on centralized data collection, making them computationally expensive and privacy-invasive. We introduce PluralLLM a federated learning-based approach that enables multiple user groups to collaboratively train a transformer-based preference predictor without sharing sensitive data, which can also serve as a reward model for aligning LLMs. Our method leverages Federated Averaging (FedAvg) to aggregate preference updates efficiently, achieving 46% faster convergence, a 4% improvement in alignment scores, and nearly the same group fairness measure as in centralized training. Evaluated on a Q/A preference alignment task, PluralLLM demonstrates that federated preference learning offers a scalable and privacy-preserving alternative for aligning LLMs with diverse human values.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
Srewa, Mahmoud
Zhao, Tianyu
Elmalaki, Salma
Machine Learning
Computation and Language
Ensuring Large Language Models (LLMs) align with diverse human preferences while preserving privacy and fairness remains a challenge. Existing methods, such as Reinforcement Learning from Human Feedback (RLHF), rely on centralized data collection, making them computationally expensive and privacy-invasive. We introduce PluralLLM a federated learning-based approach that enables multiple user groups to collaboratively train a transformer-based preference predictor without sharing sensitive data, which can also serve as a reward model for aligning LLMs. Our method leverages Federated Averaging (FedAvg) to aggregate preference updates efficiently, achieving 46% faster convergence, a 4% improvement in alignment scores, and nearly the same group fairness measure as in centralized training. Evaluated on a Q/A preference alignment task, PluralLLM demonstrates that federated preference learning offers a scalable and privacy-preserving alternative for aligning LLMs with diverse human values.
title PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.09925