Offline Clustering of Preference Learning with Active-data Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jingyuan, Ghaffari, Fatemeh, Wang, Xuchuang, Liu, Xutong, Hajiesmaili, Mohammad, Joe-Wong, Carlee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909879499751424
author Liu, Jingyuan
Ghaffari, Fatemeh
Wang, Xuchuang
Liu, Xutong
Hajiesmaili, Mohammad
Joe-Wong, Carlee
author_facet Liu, Jingyuan
Ghaffari, Fatemeh
Wang, Xuchuang
Liu, Xutong
Hajiesmaili, Mohammad
Joe-Wong, Carlee
contents Preference learning from pairwise feedback is a widely adopted framework in applications such as reinforcement learning with human feedback and recommendations. In many practical settings, however, user interactions are limited or costly, making offline preference learning necessary. Moreover, real-world preference learning often involves users with different preferences. For example, annotators from different backgrounds may rank the same responses differently. This setting presents two central challenges: (1) identifying similarity across users to effectively aggregate data, especially under scenarios where offline data is imbalanced across dimensions, and (2) handling the imbalanced offline data where some preference dimensions are underrepresented. To address these challenges, we study the Offline Clustering of Preference Learning problem, where the learner has access to fixed datasets from multiple users with potentially different preferences and aims to maximize utility for a test user. To tackle the first challenge, we first propose Off-C$^2$PL for the pure offline setting, where the learner relies solely on offline data. Our theoretical analysis provides a suboptimality bound that explicitly captures the tradeoff between sample noise and bias. To address the second challenge of inbalanced data, we extend our framework to the setting with active-data augmentation where the learner is allowed to select a limited number of additional active-data for the test user based on the cluster structure learned by Off-C$^2$PL. In this setting, our second algorithm, A$^2$-Off-C$^2$PL, actively selects samples that target the least-informative dimensions of the test user's preference. We prove that these actively collected samples contribute more effectively than offline ones. Finally, we validate our theoretical results through simulations on synthetic and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline Clustering of Preference Learning with Active-data Augmentation
Liu, Jingyuan
Ghaffari, Fatemeh
Wang, Xuchuang
Liu, Xutong
Hajiesmaili, Mohammad
Joe-Wong, Carlee
Machine Learning
Preference learning from pairwise feedback is a widely adopted framework in applications such as reinforcement learning with human feedback and recommendations. In many practical settings, however, user interactions are limited or costly, making offline preference learning necessary. Moreover, real-world preference learning often involves users with different preferences. For example, annotators from different backgrounds may rank the same responses differently. This setting presents two central challenges: (1) identifying similarity across users to effectively aggregate data, especially under scenarios where offline data is imbalanced across dimensions, and (2) handling the imbalanced offline data where some preference dimensions are underrepresented. To address these challenges, we study the Offline Clustering of Preference Learning problem, where the learner has access to fixed datasets from multiple users with potentially different preferences and aims to maximize utility for a test user. To tackle the first challenge, we first propose Off-C$^2$PL for the pure offline setting, where the learner relies solely on offline data. Our theoretical analysis provides a suboptimality bound that explicitly captures the tradeoff between sample noise and bias. To address the second challenge of inbalanced data, we extend our framework to the setting with active-data augmentation where the learner is allowed to select a limited number of additional active-data for the test user based on the cluster structure learned by Off-C$^2$PL. In this setting, our second algorithm, A$^2$-Off-C$^2$PL, actively selects samples that target the least-informative dimensions of the test user's preference. We prove that these actively collected samples contribute more effectively than offline ones. Finally, we validate our theoretical results through simulations on synthetic and real-world datasets.
title Offline Clustering of Preference Learning with Active-data Augmentation
topic Machine Learning
url https://arxiv.org/abs/2510.26301