A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xingyu, Wu, Yulian, Orabona, Francesco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908373503442944
author Zhou, Xingyu
Wu, Yulian
Orabona, Francesco
author_facet Zhou, Xingyu
Wu, Yulian
Orabona, Francesco
contents In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corruption. Specifically, under linear modeling assumptions, we present a unified analysis covering both reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) under different privacy-corruption scenarios, such as Local differential privacy-then-Corruption (LTC), where human preference labels are privatized before being corrupted by an adversary, and Corruption-then-Local differential privacy (CTL), where labels are corrupted before privacy protection. Our analysis leverages a reduction framework that reduces the offline alignment problem under linear modeling assumptions to parameter estimation in logistic regression. This framework allows us to establish an interesting separation result between LTC and CTL, demonstrating that LTC presents a greater challenge than CTL in offline alignment, even under linear models. As important by-products, our findings also advance the state-of-the-art theoretical results in offline alignment under privacy-only or corruption-only scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
Zhou, Xingyu
Wu, Yulian
Orabona, Francesco
Machine Learning
Artificial Intelligence
In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corruption. Specifically, under linear modeling assumptions, we present a unified analysis covering both reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) under different privacy-corruption scenarios, such as Local differential privacy-then-Corruption (LTC), where human preference labels are privatized before being corrupted by an adversary, and Corruption-then-Local differential privacy (CTL), where labels are corrupted before privacy protection. Our analysis leverages a reduction framework that reduces the offline alignment problem under linear modeling assumptions to parameter estimation in logistic regression. This framework allows us to establish an interesting separation result between LTC and CTL, demonstrating that LTC presents a greater challenge than CTL in offline alignment, even under linear models. As important by-products, our findings also advance the state-of-the-art theoretical results in offline alignment under privacy-only or corruption-only scenarios.
title A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.15694