Optimizing Language Models for Human Preferences is a Causal Inference Problem

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Victoria, Ben-Michael, Eli, Morency, Louis-Philippe
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911907379675136
author Lin, Victoria
Ben-Michael, Eli
Morency, Louis-Philippe
author_facet Lin, Victoria
Ben-Michael, Eli
Morency, Louis-Philippe
contents As large language models (LLMs) see greater use in academic and commercial settings, there is increasing interest in methods that allow language models to generate texts aligned with human preferences. In this paper, we present an initial exploration of language model optimization for human preferences from direct outcome datasets, where each sample consists of a text and an associated numerical outcome measuring the reader's response. We first propose that language model optimization should be viewed as a causal problem to ensure that the model correctly learns the relationship between the text and the outcome. We formalize this causal language optimization problem, and we develop a method--causal preference optimization (CPO)--that solves an unbiased surrogate objective for the problem. We further extend CPO with doubly robust CPO (DR-CPO), which reduces the variance of the surrogate objective while retaining provably strong guarantees on bias. Finally, we empirically demonstrate the effectiveness of (DR-)CPO in optimizing state-of-the-art LLMs for human preferences on direct outcome data, and we validate the robustness of DR-CPO under difficult confounding conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14979
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Language Models for Human Preferences is a Causal Inference Problem
Lin, Victoria
Ben-Michael, Eli
Morency, Louis-Philippe
Machine Learning
Computation and Language
Methodology
As large language models (LLMs) see greater use in academic and commercial settings, there is increasing interest in methods that allow language models to generate texts aligned with human preferences. In this paper, we present an initial exploration of language model optimization for human preferences from direct outcome datasets, where each sample consists of a text and an associated numerical outcome measuring the reader's response. We first propose that language model optimization should be viewed as a causal problem to ensure that the model correctly learns the relationship between the text and the outcome. We formalize this causal language optimization problem, and we develop a method--causal preference optimization (CPO)--that solves an unbiased surrogate objective for the problem. We further extend CPO with doubly robust CPO (DR-CPO), which reduces the variance of the surrogate objective while retaining provably strong guarantees on bias. Finally, we empirically demonstrate the effectiveness of (DR-)CPO in optimizing state-of-the-art LLMs for human preferences on direct outcome data, and we validate the robustness of DR-CPO under difficult confounding conditions.
title Optimizing Language Models for Human Preferences is a Causal Inference Problem
topic Machine Learning
Computation and Language
Methodology
url https://arxiv.org/abs/2402.14979