One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Hongru, Li, Yongqi, Yu, Tiezheng, Zhu, Fengbin, Wang, Wenjie, Feng, Fuli, Li, Wenjie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908979451396096
author Cai, Hongru
Li, Yongqi
Yu, Tiezheng
Zhu, Fengbin
Wang, Wenjie
Feng, Fuli
Li, Wenjie
author_facet Cai, Hongru
Li, Yongqi
Yu, Tiezheng
Zhu, Fengbin
Wang, Wenjie
Feng, Fuli
Li, Wenjie
contents Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individual users. This relies on personalized reward models that capture user-specific preferences and automatically provide individualized feedback. However, developing these models faces two critical challenges: the scarcity of feedback from individual users and the need for efficient adaptation to unseen users. We argue that addressing these constraints requires a paradigm shift from fitting data to learn user preferences to learn the process of preference adaptation. To realize this, we propose Meta Reward Modeling (MRM), which reformulates personalized reward modeling as a meta-learning problem. Specifically, we represent each user's reward model as a weighted combination of base reward functions, and optimize the initialization of these weights using a Model-Agnostic Meta-Learning (MAML)-style framework to support fast adaptation under limited feedback. To ensure robustness, we introduce the Robust Personalization Objective (RPO), which places greater emphasis on hard-to-learn users during meta optimization. Extensive experiments on personalized preference datasets validate that MRM enhances few-shot personalization, improves user robustness, and consistently outperforms baselines. We release code at https://github.com/ModalityDance/MRM.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18731
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
Cai, Hongru
Li, Yongqi
Yu, Tiezheng
Zhu, Fengbin
Wang, Wenjie
Feng, Fuli
Li, Wenjie
Computation and Language
Artificial Intelligence
Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individual users. This relies on personalized reward models that capture user-specific preferences and automatically provide individualized feedback. However, developing these models faces two critical challenges: the scarcity of feedback from individual users and the need for efficient adaptation to unseen users. We argue that addressing these constraints requires a paradigm shift from fitting data to learn user preferences to learn the process of preference adaptation. To realize this, we propose Meta Reward Modeling (MRM), which reformulates personalized reward modeling as a meta-learning problem. Specifically, we represent each user's reward model as a weighted combination of base reward functions, and optimize the initialization of these weights using a Model-Agnostic Meta-Learning (MAML)-style framework to support fast adaptation under limited feedback. To ensure robustness, we introduce the Robust Personalization Objective (RPO), which places greater emphasis on hard-to-learn users during meta optimization. Extensive experiments on personalized preference datasets validate that MRM enhances few-shot personalization, improves user robustness, and consistently outperforms baselines. We release code at https://github.com/ModalityDance/MRM.
title One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.18731