Beyond RLHF: A Unified Theoretical Framework of Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yun, Jihun, Kim, Juno, Park, Jongho, Kim, Junhyuck, Ryu, Jongha Jon, Cho, Jaewoong, Jun, Kwang-Sung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916021577711616
author Yun, Jihun
Kim, Juno
Park, Jongho
Kim, Junhyuck
Ryu, Jongha Jon
Cho, Jaewoong
Jun, Kwang-Sung
author_facet Yun, Jihun
Kim, Juno
Park, Jongho
Kim, Junhyuck
Ryu, Jongha Jon
Cho, Jaewoong
Jun, Kwang-Sung
contents Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However, existing theories do not provide strong justification for the RLHF objective itself and do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework for alignment, we ask under what assumptions can we derive existing or new training objectives and obtain theoretical guarantees. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to propose three principled alignment objectives: preference maximum likelihood estimation, preference distillation, and reverse KL minimization. We prove that they all enjoy strong non-asymptotic $O(1/n)$ convergence to the target LM, naturally avoiding degeneracy. In particular, reverse KL highly resembles the RLHF objective, providing strong justification for RLHF. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Finally, empirical results indicate that the proposed objectives are competitive with strong baselines across several tasks and models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond RLHF: A Unified Theoretical Framework of Alignment
Yun, Jihun
Kim, Juno
Park, Jongho
Kim, Junhyuck
Ryu, Jongha Jon
Cho, Jaewoong
Jun, Kwang-Sung
Machine Learning
Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However, existing theories do not provide strong justification for the RLHF objective itself and do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework for alignment, we ask under what assumptions can we derive existing or new training objectives and obtain theoretical guarantees. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to propose three principled alignment objectives: preference maximum likelihood estimation, preference distillation, and reverse KL minimization. We prove that they all enjoy strong non-asymptotic $O(1/n)$ convergence to the target LM, naturally avoiding degeneracy. In particular, reverse KL highly resembles the RLHF objective, providing strong justification for RLHF. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Finally, empirical results indicate that the proposed objectives are competitive with strong baselines across several tasks and models.
title Beyond RLHF: A Unified Theoretical Framework of Alignment
topic Machine Learning
url https://arxiv.org/abs/2506.01523