Preference Learning for AI Alignment: a Causal Perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kobalczyk, Katarzyna, van der Schaar, Mihaela
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914549925412864
author Kobalczyk, Katarzyna
van der Schaar, Mihaela
author_facet Kobalczyk, Katarzyna
van der Schaar, Mihaela
contents Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs. In this work, we propose to frame this problem in a causal paradigm, providing the rich toolbox of causality to identify the persistent challenges, such as causal misidentification, preference heterogeneity, and confounding due to user-specific factors. Inheriting from the literature of causal inference, we identify key assumptions necessary for reliable generalisation and contrast them with common data collection practices. We illustrate failure modes of naive reward models and demonstrate how causally-inspired approaches can improve model robustness. Finally, we outline desiderata for future research and practices, advocating targeted interventions to address inherent limitations of observational data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preference Learning for AI Alignment: a Causal Perspective
Kobalczyk, Katarzyna
van der Schaar, Mihaela
Artificial Intelligence
Machine Learning
Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs. In this work, we propose to frame this problem in a causal paradigm, providing the rich toolbox of causality to identify the persistent challenges, such as causal misidentification, preference heterogeneity, and confounding due to user-specific factors. Inheriting from the literature of causal inference, we identify key assumptions necessary for reliable generalisation and contrast them with common data collection practices. We illustrate failure modes of naive reward models and demonstrate how causally-inspired approaches can improve model robustness. Finally, we outline desiderata for future research and practices, advocating targeted interventions to address inherent limitations of observational data.
title Preference Learning for AI Alignment: a Causal Perspective
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.05967