Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhaowei, Bai, Fengshuo, Chen, Qizhi, Ma, Chengdong, Wang, Mingzhi, Sun, Haoran, Zheng, Zilong, Yang, Yaodong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913889883521024
author Zhang, Zhaowei
Bai, Fengshuo
Chen, Qizhi
Ma, Chengdong
Wang, Mingzhi
Sun, Haoran
Zheng, Zilong
Yang, Yaodong
author_facet Zhang, Zhaowei
Bai, Fengshuo
Chen, Qizhi
Ma, Chengdong
Wang, Mingzhi
Sun, Haoran
Zheng, Zilong
Yang, Yaodong
contents How to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, changing, and diverse regarding culture, values, or time. This leads to the problem that the actual user preferences often do not coincide with those trained by the model developers in the practical use of LLMs. Since we cannot collect enough data and retrain for every demand, researching efficient real-time preference adaptation methods based on the backbone LLMs during test time is important. To this end, we introduce Amulet, a novel, training-free framework that formulates the decoding process of every token as a separate online learning problem with the guidance of simple user-provided prompts, thus enabling real-time optimization to satisfy users' personalized preferences. To reduce the computational cost brought by this optimization process for each token, we additionally provide a closed-form solution for each iteration step of the optimization process, thereby reducing the computational time cost to a negligible level. The detailed experimental results demonstrate that Amulet can achieve significant performance improvements in rich settings with combinations of different LLMs, datasets, and user preferences, while maintaining acceptable computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
Zhang, Zhaowei
Bai, Fengshuo
Chen, Qizhi
Ma, Chengdong
Wang, Mingzhi
Sun, Haoran
Zheng, Zilong
Yang, Yaodong
Computation and Language
Machine Learning
I.2.7
How to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, changing, and diverse regarding culture, values, or time. This leads to the problem that the actual user preferences often do not coincide with those trained by the model developers in the practical use of LLMs. Since we cannot collect enough data and retrain for every demand, researching efficient real-time preference adaptation methods based on the backbone LLMs during test time is important. To this end, we introduce Amulet, a novel, training-free framework that formulates the decoding process of every token as a separate online learning problem with the guidance of simple user-provided prompts, thus enabling real-time optimization to satisfy users' personalized preferences. To reduce the computational cost brought by this optimization process for each token, we additionally provide a closed-form solution for each iteration step of the optimization process, thereby reducing the computational time cost to a negligible level. The detailed experimental results demonstrate that Amulet can achieve significant performance improvements in rich settings with combinations of different LLMs, datasets, and user preferences, while maintaining acceptable computational efficiency.
title Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
topic Computation and Language
Machine Learning
I.2.7
url https://arxiv.org/abs/2502.19148