Pluralistic Off-policy Evaluation and Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chengkai, Wu, Junda, Xie, Zhouhang, Xia, Yu, Wang, Rui, Yu, Tong, Mitra, Subrata, McAuley, Julian, Yao, Lina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918146569404416
author Huang, Chengkai
Wu, Junda
Xie, Zhouhang
Xia, Yu
Wang, Rui
Yu, Tong
Mitra, Subrata
McAuley, Julian
Yao, Lina
author_facet Huang, Chengkai
Wu, Junda
Xie, Zhouhang
Xia, Yu
Wang, Rui
Yu, Tong
Mitra, Subrata
McAuley, Julian
Yao, Lina
contents Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that differ substantially from the evaluated LLMs, and existing off-policy estimators focus solely on overall utility while ignoring preference pluralism. Extending Off-Policy Evaluation (OPE) to pluralistic preference alignment, therefore, remains an open question. Thus, we propose the Pluralistic Off-Policy Evaluation (POPE), the first framework for offline pluralistic preference evaluation and alignment in LLMs. POPE includes a unified reward function that combines (1) a collaborative utility component derived from human preference signals (e.g., upvotes or relevance scores) and (2) a diversity component inspired by entropy-based coverage measures, together reflecting pluralistic alignment. Furthermore, to estimate this reward from logged interactions, we derive decomposable inverse propensity scoring (IPS) estimators that separately evaluate relevance and diversity. Theoretically, we prove that our decomposed IPS estimators establish a lower bound on their variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance pluralistic alignment. Empirical results demonstrate that POPE efficiently enhances pluralistic response generation and maintains the models' general capabilities on downstream tasks
format Preprint
id arxiv_https___arxiv_org_abs_2509_19333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pluralistic Off-policy Evaluation and Alignment
Huang, Chengkai
Wu, Junda
Xie, Zhouhang
Xia, Yu
Wang, Rui
Yu, Tong
Mitra, Subrata
McAuley, Julian
Yao, Lina
Computation and Language
Artificial Intelligence
Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that differ substantially from the evaluated LLMs, and existing off-policy estimators focus solely on overall utility while ignoring preference pluralism. Extending Off-Policy Evaluation (OPE) to pluralistic preference alignment, therefore, remains an open question. Thus, we propose the Pluralistic Off-Policy Evaluation (POPE), the first framework for offline pluralistic preference evaluation and alignment in LLMs. POPE includes a unified reward function that combines (1) a collaborative utility component derived from human preference signals (e.g., upvotes or relevance scores) and (2) a diversity component inspired by entropy-based coverage measures, together reflecting pluralistic alignment. Furthermore, to estimate this reward from logged interactions, we derive decomposable inverse propensity scoring (IPS) estimators that separately evaluate relevance and diversity. Theoretically, we prove that our decomposed IPS estimators establish a lower bound on their variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance pluralistic alignment. Empirical results demonstrate that POPE efficiently enhances pluralistic response generation and maintains the models' general capabilities on downstream tasks
title Pluralistic Off-policy Evaluation and Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.19333