MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Jingyan, Yao, Jiarui, Yang, Rui, Sun, Yifan, Luo, Feng, Pan, Rui, Zhang, Tong, Zhao, Han
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911169023836160
author Shen, Jingyan
Yao, Jiarui
Yang, Rui
Sun, Yifan
Luo, Feng
Pan, Rui
Zhang, Tong
Zhao, Han
author_facet Shen, Jingyan
Yao, Jiarui
Yang, Rui
Sun, Yifan
Luo, Feng
Pan, Rui
Zhang, Tong
Zhao, Han
contents Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences. Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment. Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error. While existing solutions, such as multi-objective learning with fine-grained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values. In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations. In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences. In the second stage, MiCRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision. Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
Shen, Jingyan
Yao, Jiarui
Yang, Rui
Sun, Yifan
Luo, Feng
Pan, Rui
Zhang, Tong
Zhao, Han
Artificial Intelligence
Computation and Language
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences. Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment. Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error. While existing solutions, such as multi-objective learning with fine-grained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values. In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations. In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences. In the second stage, MiCRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision. Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization.
title MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.24846