Towards User-level Private Reinforcement Learning with Human Feedback

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jiaming, Lei, Mingxi, Ding, Meng, Li, Mengdi, Xiang, Zihang, Xu, Difei, Xu, Jinhui, Wang, Di
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909509329354752
author Zhang, Jiaming
Lei, Mingxi
Ding, Meng
Li, Mengdi
Xiang, Zihang
Xu, Difei
Xu, Jinhui
Wang, Di
author_facet Zhang, Jiaming
Lei, Mingxi
Ding, Meng
Li, Mengdi
Xiang, Zihang
Xu, Difei
Xu, Jinhui
Wang, Di
contents Reinforcement Learning with Human Feedback (RLHF) has emerged as an influential technique, enabling the alignment of large language models (LLMs) with human preferences. Despite the promising potential of RLHF, how to protect user preference privacy has become a crucial issue. Most previous work has focused on using differential privacy (DP) to protect the privacy of individual data. However, they have concentrated primarily on item-level privacy protection and have unsatisfactory performance for user-level privacy, which is more common in RLHF. This study proposes a novel framework, AUP-RLHF, which integrates user-level label DP into RLHF. We first show that the classical random response algorithm, which achieves an acceptable performance in item-level privacy, leads to suboptimal utility when in the user-level settings. We then establish a lower bound for the user-level label DP-RLHF and develop the AUP-RLHF algorithm, which guarantees $(\varepsilon, δ)$ user-level privacy and achieves an improved estimation error. Experimental results show that AUP-RLHF outperforms existing baseline methods in sentiment generation and summarization tasks, achieving a better privacy-utility trade-off.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards User-level Private Reinforcement Learning with Human Feedback
Zhang, Jiaming
Lei, Mingxi
Ding, Meng
Li, Mengdi
Xiang, Zihang
Xu, Difei
Xu, Jinhui
Wang, Di
Machine Learning
Artificial Intelligence
Reinforcement Learning with Human Feedback (RLHF) has emerged as an influential technique, enabling the alignment of large language models (LLMs) with human preferences. Despite the promising potential of RLHF, how to protect user preference privacy has become a crucial issue. Most previous work has focused on using differential privacy (DP) to protect the privacy of individual data. However, they have concentrated primarily on item-level privacy protection and have unsatisfactory performance for user-level privacy, which is more common in RLHF. This study proposes a novel framework, AUP-RLHF, which integrates user-level label DP into RLHF. We first show that the classical random response algorithm, which achieves an acceptable performance in item-level privacy, leads to suboptimal utility when in the user-level settings. We then establish a lower bound for the user-level label DP-RLHF and develop the AUP-RLHF algorithm, which guarantees $(\varepsilon, δ)$ user-level privacy and achieves an improved estimation error. Experimental results show that AUP-RLHF outperforms existing baseline methods in sentiment generation and summarization tasks, achieving a better privacy-utility trade-off.
title Towards User-level Private Reinforcement Learning with Human Feedback
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.17515